šŸ†

Infernet Econometrics Benchmark

CamoAiLab Ā· HKU CAMO

An evaluation benchmark for large‑language‑model agents, measuring the capability to reproduce real‑world econometric empirical studies: generating executable code, replicating regression outputs, and recovering correct treatment‑effect coefficient signs.

Leaderboard

Official Benchmark
# Model Score
Loading rankings…

Leaderboard

Official Benchmark
# Model Score
Loading rankings…

Leaderboard

Official Benchmark
# Model Score
Loading rankings…

Metric Definitions

Note: Codex serves as a strong code‑specialized upper‑bound baseline with high overall metrics across all three evaluation dimensions. However, it only supports one‑shot code generation without agent‑level interactive planning or multi‑round revision capabilities, and its public API is no longer available. MetricsAI outperforms vanilla‑LLM and general‑purpose‑agent baselines for interactive real‑world econometric‑research workflows.

Dataset: CamoAiLab/Infernet