Recursive self-improvement is now something frontier labs try to measure, but what it is, and how to benchmark it, remain open questions. Whatever the definition, training an AI model end to end on real infrastructure sits at its core. This volume audits all 626 public task environments in twelve benchmarks against the lifecycle of production LLMs and maps what they actually exercise.