A product discussed on AI Engineer.

The Agent Behind the Curtain: Building the Oz Cloud Agent Platform — Safia Abdalla, Warp
Aug 22, 2026 · 20:51
Safia Abdalla, an engineer at Warp, explains how her team built the Oz cloud agent platform on the principle that platforms should absorb complexity before the user sees it. Agents run in managed or self-hosted sandboxes; the platform supports multiple harnesses consistently. Users can orchestrate sub-agents by prompt or via an API exposed across the stack, and non-engineers used the SDK to build Slack bots for triaging social mentions. Abdalla recounts open-sourcing Warp: stars grew from about 20,000 to over 60,000, with thousands of PRs; agents triage issues and review every PR, so humans only see high-signal ones. She rejects 'software factory' as losing the people in it, offering a potter's workshop: stations, sourcing, verification, and an observable, improvable, cost-effective system.

Fine tune 20 Llama Models in 5 Minutes: Santosh Radha
Feb 9, 2025 · 6:26
Santosh Radha, Head of Product/Research at Agnostiq, demonstrates Covalent, an open-source platform that lets users fine-tune and deploy hundreds of Llama models directly from Python without Kubernetes or Docker. By adding a single decorator to Python functions, users specify GPU requirements (e.g., H100 with 48 GB, 18-hour limit) and run them on remote compute, paying only for actual usage (e.g., 87 cents for 6 minutes on an L14, 11 cents on a V100). Covalent supports job submission, inference endpoints with custom autoscaling (e.g., scale to 10 GPUs at 9 AM daily), and automated workflows for training, evaluation, and deployment. Radha shows a workflow that iterates over 20 models, fine-tunes each, evaluates accuracy, sorts, and deploys the best—all from a Jupyter notebook with a single dispatch call. The talk, recorded at the AI Engineer World's Fair, emphasizes eliminating infrastructure overhead for accelerated compute.
Powered by PodHood