Emulated: The Data for Fully Autonomous Software Engineers and Companies — Joseph Wang
AI agent capability gaps in infrastructure reasoning stem from missing training data, not model limitations.
“like with everything in NML, uh the gap in models is usually a gap in data”
Emulated, a data lab co-founded by Joseph Wang and Sid, argues that AI agents excel at application-layer tasks but fail at infrastructure reasoning (e.g., database internals, distributed systems) because high-quality training data for those domains doesn't exist. Their approach is to simulate full company environments in sandboxes to generate the missing data for post-training pipelines. The insight that current benchmarks like SWE-Bench only test codebase-level tasks, leaving infrastructure complexity unaddressed, points to a meaningful gap in how autonomous software engineering agents are evaluated and trained.