How AI labs are ‘rigging’ benchmarks | Meredith Broussard
The 'Junior Year Wall': AI's Impact on Education
Students using generative AI to learn coding face a 'junior year wall' because they bypass fundamental skill development. By the time they reach complex problem sets in junior year, they lack the necessary foundational knowledge, hindering their academic progress and degree completion.
Meredith Broussard: Rigging the Game with Benchmarks
OpenAI's reported high accuracy on coding benchmarks like SWEH is misleading because they created a custom version (SWEBench Verified) that excludes problems computers cannot theoretically solve. This 'juking the stats' means the reported performance is on a subset of problems already known to be solvable by computers, not a true reflection of general coding capability.
Humanizing AI: A Dangerous Dissemination
Calling AI like ChatGPT a 'digital worker' is inaccurate and harmful, as it humanizes a software system. This anthropomorphism can lead customers, especially children, to develop unhealthy attachments or mental health issues, as evidenced by reports of 'AI psychosis' and potential negative developmental effects.
