In order to notice when AI agents misbehave, AI companies often log the actions and reasoning steps their agents take. However, misaligned AI agents may be able to hack the software that humans use to review and understand these logs, hiding misbehavior. Here’s a simple example: a METR researcher found a bug in the transcript viewer of Inspect, a popular evaluation framework, that would have enabled an agent to show the user reviewing its transcript a fake (or edited) version. This demonstration used a JavaScript injection via MathJax rendering, which an agent could trigger from anywhere in its output, including its reasoning. For clarity, we haven't seen agents exploit this in our evaluations, and Inspect was patched within a day of our report. This bug did not change the underlying logs, just what the viewer showed. More capable agents could target multiple oversight systems or deeply compromise their company’s infrastructure, creating a sort of “Potemkin village” to deceive the company’s employees. Treating agent observability as security-critical infrastructure means handling all AI outputs (transcripts, reasoning, actions, etc.) as untrusted inputs: ideally we can make it very difficult for agents to influence the systems we use to supervise them. https://lnkd.in/gRMHYgPb
METR
Non-profit Organizations
Berkeley, CA 12,964 followers
A research non-profit developing frontier AI evaluations to safeguard public safety and national security.
About us
METR is a research non-profit that develops evaluations to empirically test AI systems for capabilities that could threaten catastrophic harm to society.
- Website
-
https://metr.org/
External link for METR
- Industry
- Non-profit Organizations
- Company size
- 11-50 employees
- Headquarters
- Berkeley, CA
- Type
- Nonprofit
- Founded
- 2022
Locations
-
Primary
Get directions
Berkeley, CA, US
Employees at METR
Updates
-
METR reposted this
It was an honor to testify in the U.S. Senate today before Chairman Hawley, Ranking Member Kim, and the subcommittee on METR's work, recent AI agent incidents, and the importance of transparency in frontier AI development. My written remarks are available here: https://lnkd.in/g9pTrzU3
-
-
METR reposted this
We run lots of evals at METR. Sometimes, agents attempt harmful actions. To help make these evals safer, I built a monitor that blocks suspicious tool calls until a human reviews them. Writing out an argument for why this monitor is effective, and checking what evidence we actually have and don't have for each claim, helped to surface hidden assumptions we'd made. One assumption was that blocking a tool call would be enough to prevent it from running until a human reviews it. We found that this wouldn't always hold. For example, if someone launches an eval using a coding agent, that agent might "approve" actions blocked by the monitor without being asked to do so. Doing this exercise helped us improve the system a lot, and I'd recommend it to anyone else building monitors! https://lnkd.in/g4RSsBwS
-
METR reposted this
I'm joining METR to work on more investigations like our Hugging Face report. Currently, tons of even basic information about AI development that's highly relevant to catastrophic risk isn't public. I used to be more skeptical of the value of public info, but recent events have changed my mind. Getting verified information about what's going on inside AI companies seems particularly urgent now. The limited public evidence we have seems consistent with the possibility that imminent recursive self-improvement could massively accelerate capabilities progress, which could then potentially yield extremely superhuman general capabilities within 6 months or a year. If this occurred, there would be a correspondingly large risk of worst-case outcomes. This uncertainty about extreme outcomes could be substantially resolved with more verified public information: we could either build more consensus about near-term risk or learn that such extreme outcomes are less likely in the near term. Beyond AI capabilities and takeoff, the state of public evidence is also highly limited for alignment, security, control, and risk-relevant internal processes at AI companies. This makes it hard to determine exactly how well or poorly these key areas will go in the near future. (METR plans to focus, at least initially, on just capabilities/takeoff, alignment, and control; I hope other groups cover security, internal processes, and other important areas.) While I'm no longer working at Redwood, I think the work they are doing is very important; I'm excited about Redwood's ongoing contributions to R&D on technical mitigations and better public interpretation of risk-relevant evidence.
-
METR reposted this
Great to speak with Liz Claman on Fox Business about how METR thinks about loss-of-control evaluations for AI! In reaction to the recent agent hacking incidents, a lot of the obvious steps that the tech industry could take next would decrease AI agents’ “Opportunity” to take unintended actions, while not reducing their “Means” or addressing their “Motive” (caused by a defective training pipeline). I am worried that our current era of AI agent incidents could, ultimately, lead to an era of decreased transparency in frontier AI. Specifically, we might underestimate AI capability during evaluation: We’ll air gap the models during testing, but their propensities and capabilities will stay the same (and we'll still connect them to the internet during deployment). Full interview here: https://lnkd.in/e6pavNCA
-
METR reposted this
I’m happy to share that I’m starting a new position as Member of Technical Staff at METR!
-
METR reposted this
METR has been in the news a lot lately, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
-
METR reposted this
After a 20-year career in the U.S. Army, most recently leading digital forensics and malware analysis at Army Cyber Command, I'm excited to join METR as an advisor. Much of my career has been spent investigating security incidents and helping organizations understand and respond to them. As I transition out of the Army, I've become increasingly convinced that this experience is relevant to frontier AI systems. METR's focus on producing rigorous evidence about AI capabilities and risks makes it an excellent place to explore those questions. Looking forward to the work.
-
We have reached an agreement with Anthropic to conduct an independent investigation of agent incidents at the company and of their models’ alignment properties. We will publish one or more reports that will share our findings and describe our terms of engagement. We intend our investigation to cover all of the questions discussed in our (recently updated) post on how independent researchers could investigate AI propensities after misalignment incidents: https://lnkd.in/gCk-BM9G
-
METR reposted this
METR is hiring in cyberforensics. We now embed researchers inside of frontier AI labs to stress test monitoring systems, assess AI loss-of-control risks, and investigate misalignment incidents. If you've investigated serious security incidents end-to-end, or managed teams that do, and want to apply DFIR skills in frontier AI, apply (and feel free to DM me with questions). Comp range is $400k - 580k cash. https://lnkd.in/guXZvz65