AI Researcher Lun Wang Departs DeepMind, Spotlights Gaps in LLM Evaluation

AI Researcher Lun Wang Departs DeepMind, Spotlights Gaps in LLM Evaluation

Artificial intelligence keeps getting smarter. What systems are used to measure that intelligence? Not so much.

That’s the warning from Lun Wang, a senior researcher who recently left Google DeepMind and used his departure to spotlight what he sees as one of the biggest blind spots in modern AI development: the way large language models are evaluated.

In a post shared on X, formerly Twitter, Wang argued that current AI benchmarks are no longer sufficient for measuring increasingly advanced systems. His proposed solution — “self-evolving evals” — could reshape how the industry tests safety, intelligence, and reliability in the next generation of AI models.

The concern lands at a critical moment for the AI industry. Companies are racing to build more capable systems, but researchers increasingly worry that existing evaluation methods are too static to keep up.

TL;DR

Why AI Evaluation Has Become a Major Problem

AI benchmarks once served a simple purpose: to compare one model against another.

Researchers would feed systems a standardized set of questions, coding tasks, or reasoning problems. Scores helped determine which models performed better. But as AI capabilities accelerated, those tests began losing their usefulness.

Today’s large language models can memorise benchmark datasets, exploit patterns in evaluation methods, or perform impressively in narrow tests while failing badly in real-world scenarios.

That creates a dangerous gap between what AI appears capable of and what it can actually do.

Wang described this mismatch as “the most important unsolved problem” in understanding LLMs. The statement reflects a growing concern among AI researchers that the industry’s measurement tools are lagging behind the technology itself.

The benchmark problem in plain English

Imagine giving students the same exam every year.

Eventually, students memorise the answers instead of learning the subject. Their scores rise, but their understanding may not.

Researchers say something similar is happening with AI models.

Static benchmarks can become predictable. Once models are trained on enough internet data, they may effectively “see” parts of the tests beforehand. That can inflate performance scores without reflecting genuine reasoning ability.

Consider adding an infographic here comparing:

What Are ‘Self-Evolving Evals’?

Wang’s proposed solution is straightforward in concept but difficult in execution.

Instead of fixed benchmarks, AI systems would be tested using dynamic evaluations that continuously adapt as models improve.

These “self-evolving evals” would:

The goal is to create evaluation systems that evolve at nearly the same pace as the models themselves.

Why adaptive testing matters

Current AI evaluations often focus on narrow capabilities:

But advanced AI systems can display unexpected behaviors outside those controlled settings.

For example:

Adaptive evaluations could help researchers catch those issues earlier.

The Bigger Concern: AI Safety and Trust

Wang’s warning is not just about technical accuracy. It is also about governance and public trust.

If companies rely on outdated testing methods, they could make poor decisions about:

In other words, weak evaluations can create false confidence.

That concern has become increasingly important as AI companies compete to release more powerful models at a faster pace. Many labs now emphasise “frontier AI” development, systems designed to handle increasingly complex reasoning and autonomous tasks.

But measuring those systems remains difficult.

Why current benchmarks may fail

One major issue is that benchmarks often measure performance snapshots instead of long-term behavior.

A model might pass:

Yet still behave unpredictably in live environments.

Researchers sometimes refer to this as the “capability-evaluation “gap”—the difference between benchmark success and real-world reliability.

AI Researchers Are Increasingly Questioning Benchmarks

Wang is not alone in raising concerns about AI evaluation.

Across the industry, researchers have started questioning whether benchmark culture has distorted AI progress.

Some critics argue that companies optimize models specifically to score well on popular public tests. That can create leaderboard-driven development instead of genuine advances in reasoning or safety.

Others warn that many benchmarks become obsolete too quickly.

For example:

This is partly why companies have started building private evaluation systems that are harder for models to anticipate.

Still, no universal standard exists.

What Happens Next?

The idea of self-evolving evaluations is still largely conceptual. Building them would require:

It would also require cooperation across the AI industry, something that has historically been difficult in competitive technology races.

Yet the push for better evaluations is likely to intensify.

As AI systems gain stronger reasoning abilities and broader autonomy, the industry may no longer be able to rely on old-style benchmarks designed for earlier generations of models.

Wang’s departure from Google DeepMind adds extra visibility to that debate. His comments highlight a growing realization inside the AI community: building smarter models is only half the challenge.

Understanding them may be even harder.

Why This Story Matters Beyond Silicon Valley

The benchmark debate may sound technical, but it affects everyday users more than most people realise.

AI evaluations influence:

If evaluation systems fail, the consequences can spread quickly, from misinformation problems to flawed automated decision-making.

That is why researchers increasingly see evaluation not as a side task but as a core part of responsible AI development.

And according to Wang, the industry is running out of time to modernize it.

Exit mobile version