Hugging Face Journal Club: Training AI Scientists to Replicate Research
RL-trained Qwen 27B model Faraday outperforms Claude and GPT-5 on scientific replication tasks
“they can get this so-called AI scientist agent which can outperform you know much larger models such as Claude and GBD5”
Startup Inherent trained Faraday, a Qwen 27B model using RL, to replicate missing figures from ML/AI-for-science papers — a narrow but meaningful scientific reasoning benchmark. The model outperforms much larger frontier models including Claude and GPT-5 by using Codex as a tool and a rubric-based LLM judge for training rewards rather than verifiable ground-truth signals. This is a notable data point in the emerging 'AI scientist' space showing that domain-specific RL fine-tuning on smaller models can beat scale on targeted research tasks.