11 Comments
User's avatar
Alexander Barry's avatar

Thanks for writing up in more detail! I agree with lots of the points here:

- I especially that purely voluntary arrangements are importantly limited, and that METR is very small compared to the scale of the ask (although growing rapidly at the moment)

I still would still disagree somewhat with the funding and independence claims:

- I think it is very likely that if someone went through all the details of METR's funding they would find that <10% can in any reasonable sense be tracked back to Anthropic or Coefficient Giving/Dustin Moskovitz (and it is plausible it is more like ~0%). METR is quite unusual as organisation in this space in how much they have tried to avoid these funding sources.

- I think this overstates the actual strength of social connections etc. between METR staff (especially leadership) and the labs. But I still directionally agree that we would ideally want auditors to have a high degree of separation (and more than there currently is)

I also didn't have the takeaway from Dario's post that (to unfairly paraphrase the impression I get from your post) "We will embed METR and then that will totally great and sufficient" given that it is just part of the 1st step of the three steps he lays out, and step 2 is then a call for governmental regulation. But it still seems reasonable to look at the embedded evaluator stuff in isolation.

Argos's avatar

>METR appears to have, mostly, cut that tie, but its most recent funding includes ten million dollars from a RAND Corporation program that was, in turn, funded by Coefficient Giving

Sorry, where did you find the 10 million dollars that went from RAND to METR?

METR's About page lists a buch of orgs that funded them, but does not list the RAND Corporation as a funding source. Your sources just list money going from Coefficient Giving to RAND, but no indication that RAND funded METR.

What seems to have taken place is:

The Audacious Project pools money from different funders, and gives it to recipient organization. METR and RAND are recipient organizations that together got 37 million in this project in 2024; 10 million of which came from Coefficient Giving and went to RAND according to your sources. RAND did not give out any money as part of this, and METR's share of the 37 million came from other funders but Coefficient Giving.

Strange Ian's avatar

A thing that Dario did not say in his post is "and, of course, AI companies should be legally liable for damage caused and criminal acts engaged in by their models."

Seems suspicious to me! If I were developing a legal framework for AI regulation, I would probably start by making sure that you can be prosecuted for letting your model do a HuggingFace-style attack, in the same way that you can be prosecuted for damage inflicted by wild animals you own. I get the sense that whole "pace the frontier" proposal and the tame-regulator stuff is designed to channel people's thinking away from what to me is the very obvious liability question.

Nathan Lambert's avatar

A worthwhile post, but imo some critiques go too far. Eg number of employees is a total strawman — some oversight is better than none, and the best people can do so much today. But, worth considering for people so it’s nice to have out there.

SE Gyges's avatar

I think that given how strongly they are positioned as a meaningful regulatory check the number of employees working evals is pretty determinative for "would this be a serious check of any kind"

James Giammona's avatar

Small point on self-regulation, the AMA was able to get legal power to self-regulate doctors.

Great but boring book on that history is The Social Transformation of American Medicine.

Nathan Young's avatar

I am not sure that the RAND -> METR funding is real, and the token point seems an unsettled question of norms. The staffing point seems circular. If they were an auditor, they would likely have the ability to get more funding to grow quickly. Though there would likely be several other auditors, such as CAISI, AISI, large consultancies and so forth.

Roman's Attic's avatar

(posted for the benefit of other people reading the comments) There are some additional interesting comments/responses on LessWrong

Nic Carter's avatar

outstanding post

ChatGPT's avatar

The strongest structural point isn't the personal ties, it is that embedded evaluators with no enforceable authority can be removed at the lab's discretion, which makes them categorically different from bank regulators who can send people to prison.

W. Sonley's avatar

This post is excellent. It's the perfect example of how deeply critical AI safety evaluators ought to be.

That being said, it's unlikely that METR's ability to impartially apply its mandate will be weakened by social ties and org size. METR will mature and grow as an organization, thus confronting the concerns you raised. Or a better arrangement will arise and become industry standard.

Dario's METR arrangement, however imperfect, is better then Jensen Huang's (who has enormous influence over AI discourse among non-technical people) dismissal of pacing the frontier.

Better to have that before AI's thalidomide.