Skip to content
daily-hour-news·

🔬Agent Skills Lift MLE-Bench Scores 134% at Fixed Compute

TL;DR

DisCo distills 1,000 widely used ML repos into 5,000-plus verified skills, then hands them to a research agent with backbone and budget unchanged. MLE-bench scores rise 134.3%, PaperBench 34.4%, FrontierCS 9.2% and PassNet 14.0%.

DisCo distills 1,000 widely used ML repos into 5,000-plus verified skills, then hands them to a research agent with backbone and budget unchanged. MLE-bench scores rise 134.3%, PaperBench 34.4%, FrontierCS 9.2% and PassNet 14.0%.

Agent Skills Lift MLE-Bench Scores 134% at Fixed Compute — daily-hour-news

Key Points

1

arXiv 2609.02749, submitted Sep 2 by Jianlyu Chen and 10 co-authors

2

AREX-Skill Library holds 5,000+ verified skills across 20 areas and 178 capability families

3

Backbone fixed at GPT-5.5 with the harness and execution budget held constant

4

Gains: MLE-bench +134.3%, PaperBench +34.4%, FrontierCS +9.2%, PassNet +14.0%

5

Two distillation modes: task-agnostic repo condensation and task-oriented skill synthesis

Why It Matters

The cheapest capability gain on the table right now may be packaging operational know-how as skills rather than buying a bigger model.

Quick Facts

AI agentsagent skillsMLE-benchAI researcharXivGPT-5.5

Frequently Asked Questions

Why does this matter?

The cheapest capability gain on the table right now may be packaging operational know-how as skills rather than buying a bigger model.

What happened?

DisCo distills 1,000 widely used ML repos into 5,000-plus verified skills, then hands them to a research agent with backbone and budget unchanged. MLE-bench scores rise 134.3%, PaperBench 34.4%, FrontierCS 9.2% and PassNet 14.0%.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,463 builders reading daily.

Also get