📊 Full opportunity report: Engineering Is Automated. Research Is the Residual. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Recent evidence indicates AI has achieved near-complete automation of engineering tasks in AI research. However, research activities still require human insight. This shift could transform AI development processes but leaves open questions about the limits of automation in research.
Recent empirical evidence indicates that AI systems can now automate the majority of engineering tasks involved in AI research, while research activities themselves remain less automated.
Multiple benchmarks, including CORE-Bench and MLE-Bench, show rapid progress in automating core AI engineering skills. For example, CORE-Bench, which measures research reproduction, reached 95.5% success in December 2025, with one author declaring it ‘solved.’ Similarly, MLE-Bench, assessing Kaggle competition performance, hit 64.4% in February 2026, demonstrating AI’s competitiveness with mid-tier human practitioners.
These benchmarks reveal that AI can handle complex, friction-prone tasks such as reproducing research papers and competing in ML challenges at levels approaching human experts. Meanwhile, the development of automated kernel design and infrastructure optimization further exemplifies the transition of engineering tasks into AI capabilities.
However, Clark’s analysis suggests that research activities—such as hypothesis generation, experimental design, and creative problem-solving—are less amenable to full automation, leaving a residual gap between engineering automation and research automation.
Engineering is automated.
Research is the residual.
Six skill benchmarks. Edison’s framing. The question Clark leaves open is whether research is just engineering at scale.
Jack Clark’s Import AI #455 catalogs six benchmarks measuring AI capability on AI R&D tasks and concludes “AI can today automate vast swatches, perhaps the entirety, of AI engineering.” The residual question is research. The structural read on the residual: it may not be a permanent moat.
Six skills. One trajectory.
Clark catalogs six benchmarks measuring AI capability on AI R&D-relevant tasks. Each individual benchmark could be noise. Six benchmarks moving together is a curve. The pattern is the cascade observed across the broader Clark series — visible here in the specific R&D-skill domain.

AI-Powered Shopify Store Starter Guide: A Beginner-Focused Playbook to Launch a Profitable Shopify Store Using AI Tools for Product Research, Branding, Listings, and Sales (AI-POWERED E-COMMERCE 2)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three data points. Mixed signal.
Clark provides three data points on the creative-spark question. Yes-evidence: Erdős-1051, centaur math discovery, sporadic Move-37-style moments. No-evidence: low yield, framing dependence, absence of acceleration. The mixed signal is the honest read.
The data supports two readings. Pessimistic: rare moments suggest creative insight is qualitatively distinct from engineering work. Optimistic: rare moments are an artifact of low-volume exploration; more shots on goal yields more discoveries. Both readings are consistent with Clark’s “vast swatches, perhaps the entirety” claim. They differ on the residual.

Automated Machine Learning Systems: Integrating Automated Machine Learning into Modern Data Platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five dimensions Clark gestures at but leaves underdeveloped.
Clark’s section is rigorous on the empirical evidence. Five strategic dimensions matter for the institutional response that the Clark series synthesis argues is structurally inadequate.
research paper reproduction software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Two readings. Different equilibria.
The structural question Clark leaves open: is research a permanent moat that bounds automated AI R&D, or is it engineering at scale that dissolves with more shots on goal? Both readings are consistent with the current data. They differ by orders of magnitude in consequences.
Productivity multiplier years
Recursive loop operational

Evals for AI Engineers: Systematically Measuring and Improving AI Applications
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five audiences. Asymmetric cost of being wrong.
The institutional response should not bet on inspiration being a permanent moat. If the distinction holds, capacity built is still useful. If it closes, capacity is necessary. Asymmetric cost-of-being-wrong points toward building now.
IN INDUSTRY
IN ACADEMIA
POLICYMAKERS
INVESTORS
EVERYONE ELSE
Engineering is automated. The residual is the question. The institutional response should not bet on inspiration being a permanent moat.
Implications of AI-Driven Engineering Automation
The automation of engineering tasks in AI research could dramatically accelerate development cycles, reduce costs, and shift the role of human researchers toward higher-level creative and strategic activities. This shift may challenge traditional research workflows and institutional structures, prompting a reevaluation of how AI innovation is managed and funded.
However, the persistence of research as a residual activity underscores ongoing uncertainties about the limits of automation, especially regarding hypothesis formulation, experimental insight, and scientific creativity. The potential for AI to fully automate research remains an open question, with significant implications for the future of scientific discovery and AI development.
Recent Progress in AI Engineering Capabilities
Over the past 18 months, multiple benchmarks have shown rapid improvements in AI systems’ ability to perform core engineering tasks. The CORE-Bench, measuring research reproduction, advanced from 21.5% in September 2024 to 95.5% in December 2025. Similarly, the MLE-Bench, evaluating Kaggle competition performance, improved from 16.9% in October 2024 to 64.4% in February 2026. Concurrently, advances in kernel design, including automated GPU kernel generation and infrastructure optimization, demonstrate AI’s transition from experimental to production-grade capabilities.
This pattern reflects a broader trend of AI systems approaching or surpassing human-level performance in specific engineering domains, suggesting that much of the ‘perspiration’—the routine, friction-prone work—is now automatable, leaving the more creative and hypothesis-driven research activities as the remaining challenge.
“The evidence suggests that AI can today automate vast swaths, perhaps the entirety, of AI engineering. The residual research—those parts involving hypothesis, creativity, and experimental insight—remains less automatable.”
— Thorsten Meyer
Unresolved Questions About Research Automation Limits
While engineering tasks are increasingly automated, it remains unclear how much of the research process—such as hypothesis generation, experimental design, and creative problem-solving—can be automated. The structural question Clark leaves open is whether research itself is just scaled engineering or involves inherently human creative insight that resists automation.
Further developments and empirical evidence are needed to determine whether AI can fully automate research activities or if residual human oversight will remain necessary.
Next Steps in Monitoring AI Research Automation Progress
Researchers and industry observers will continue to track benchmark progress and new developments in automated kernel design, research reproduction, and competitive ML tasks. Key milestones include further improvements in automation accuracy, the development of new benchmarks to measure research creativity, and empirical testing of AI’s capacity to generate novel hypotheses and experimental protocols. Additionally, institutional responses and policy discussions will shape how this technological shift influences AI research practices.
Key Questions
What specific tasks in AI research are now automated?
Tasks such as reproducing research papers, running experiments, and optimizing infrastructure are now largely automated, as demonstrated by benchmarks like CORE-Bench and MLE-Bench.
Does this mean AI can replace human researchers completely?
While many engineering tasks are automatable, the ability of AI to fully replace human researchers—particularly in hypothesis generation and creative problem-solving—remains uncertain and is an active area of investigation.
What are the implications for scientific progress?
If AI can automate engineering at scale, it could accelerate development cycles and reduce costs. However, the residual research activities may still require human insight, shaping future collaboration models between humans and AI.
Are there risks associated with AI automating research and engineering?
Potential risks include over-reliance on automated systems, reduced human oversight, and challenges in ensuring the quality and safety of AI-generated research. These concerns are subjects of ongoing discussion.
Source: ThorstenMeyerAI.com