Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI
Jul 31, 2026 · 19:05
DatologyAI CEO Ari Morcos argues data quality is the compute multiplier: better data steepens scaling, so the same compute buys better models. His oil-refinery approach—clean, curate, create, compose—uses synthetic rephrasing for diversity; curation let a VLM beat the public Pareto frontier with 145x less training compute and match Qwen 3.5 with 35x fewer flops per correct answer. Curating English also boosts non-English via cross-lingual transfer. For Thomson Reuters, mid-training on curated legal data lifted LegalBench 5 points without catastrophic forgetting and tripled post-training gains. Arcee's Trinity Large, trained on 17 trillion curated tokens, matched GLM-5 and Kimi and beat Claude on some tasks for under $20 million, proving data curation is cheaper than compute.