Page 3 of 23

How to Kill the Code Review — Ankit Jain, Aviator
Aug 17, 2026 · 16:26
Aviator co-founder Ankit Jain argues code review isn't just about catching bugs: it carries knowledge sharing and mentorship, and that half must survive even as AI writes and reviews code. He says reading line by line is already over — over 30% of changes merge without review. He calls spec-driven development 1970s waterfall because intent lives in the prompts, which teams discard when the PR opens. His fix: capture agent sessions, turn decisions into acceptance criteria, pair them with an 'AI slop registry' of recurring review comments, and generate test plans a verification system runs against a live preview. Reviewers review intent and evidence, not the diff. Homework: mine your last 1,000 review comments; he pitches Aviator's new Verify pilot.

Security Firewall for Agents — Ryan Dahl, Deno
Aug 17, 2026 · 19:06
Ryan Dahl, CEO of Deno, argues agents must be treated as untrusted software and introduces Claw Patrol, an MIT-licensed proxy that parses every byte leaving an agent below the HTTP layer. At Deno Deploy, agents with write access to Postgres, Kubernetes, ClickHouse, and AWS can be prompt-injected through the support system, so Opus refusing to delete the users table is not enough. Claw Patrol blocks destructive actions even when an agent spawns psql through an EKS endpoint, using HCL rules checked into Git, holds credentials so agents never see them, and can route actions to an LLM judge or Slack approval. A demo shows Codex in yolo mode trying to delete the users table and being blocked. It also includes a unit test system with fixture requests to ensure rules work.

Context Engineering in 2026 — Louis-François Bouchard, Omar Solano & Samridhi Vaid, Towards AI
Aug 17, 2026 · 1:03:26
Louis-François Bouchard, Omar Solano, and Samridhi Vaid of Towards AI test context engineering on their AI tutor: keeping full history beat every compaction technique on recall, cost, and latency because 97% of tokens were served from cache (up to 50x cheaper), making summarization a trap unless it shrinks context by more than 50x. Full history recovered specific details 95% vs 32% after summarizing, and distinctive facts survived 800k tokens. Local hardware changes it: a 32k window can't keep everything, and larger models don't widen context. Dense retrieval hit 0% recall at 400k tokens where BM25 got 100%, so they use hybrid retrieval. Rule: name the constraint before compacting.

How Web Data Infrastructure Powers the Next Generation of AI — Patricija Žemaitytė, Oxylabs
Aug 14, 2026 · 19:03
Oxylabs' Patricija Žemaitytė argues infrastructure, not better models, will power the next generation of AI, citing three client-driven projects. A video API built on a two-week deadline and a 5-petabytes-per-month floor expanded into transcripts, subtitles, search, and metadata; client had 30 petabytes and still hadn't paid. A search API rebuilt for subsecond latency and zero data retention got blocked live on a demo call, then reached 550 milliseconds average from a 4-second baseline by trimming layouts, parsers, sessions, and proxies. Scaling the web unblocker from 10,000 to 60,000 requests per second stalled at 20,000 in load testing because telemetry became part of the load, and the follow-up is already Project 150. The lesson: this is not a build-once business but an adapt-forever one.

The Rise of CaaS: Context-as-a-Service for Agentic AI — Omer Primor, Bright Data
Aug 14, 2026 · 22:20
Omer Primor of Bright Data argues that rented AI search and context-as-a-service (CaaS) lose to a self-built pipeline once query volume passes a tipping point. The web is context, not data, and it decays: social content goes stale within a day, news/finance/retail within 30 days, so context is never a one-time snapshot. His test enriching 100 sponsor companies across 25 fields found general search beat dedicated CaaS vendors on coverage, because CaaS only answers from data it already holds; costs were similar, but frequency is the real killer — every repeated query costs the same when nothing changed. He built scrapers for LinkedIn, jobs, and Crunchbase in a day, priced setup at $5,000, and put the crossover just over 15,000 entities. Owned context compounds while rented decays.

From RL to IRL — Gaurav Mishra, Amazon AGI Lab
Aug 14, 2026 · 17:46
Gaurav Mishra of the Amazon AGI Lab explains why RL for computer-use agents works in games but breaks in real life, and how his team turns failures into training data. He shows early browser-training runs where an agent guesses its password and locks the account, and clicks a sponsored button styled like the submit button, landing elsewhere. The talk catalogs partial observability, irreversibility, expiring credentials, and ambiguous success, then proposes flight-school sandboxes, process reward models, calibrated confidence, and adversarial tasks. A later trajectory shows the agent recognizing the sponsored button, refusing to guess the password, and handing off to the user, echoing the talk's point that the difference between a demo and a product is what happens after the first failed click.

The Dark Arts of Web Automation: Teaching Agents to Use Websites Like Humans — Corey Gallon, Rexmore
Aug 14, 2026 · 21:38
Corey Gallon (Rexmore) argues that an agent driving Chrome via the Chrome DevTools Protocol is just a 'meat bag with a mouse'—and details his open-source Chrome Agent tool. Preparing the talk earned him an OpenAI ban threat after his agent cleared Cloudflare Turnstile, MT Capture, Lemon, and reCAPTCHA v2 with no human in the loop. He argues for CLI over MCP: both succeeded ~83% in an Arise AI study, but CLI took 7 turns and under a minute versus MCP's 71 round trips and 8 minutes, and CLI can be 75x cheaper in tokens. The method is a sense/act/verify loop up a three-rung ladder: synthetic clicks, trusted CDP input, then human mouse paths with jitter, deliberate overshoot. For reCAPTCHA, deterministic code drives and rearms each round while the agent names grid tiles, as speed beats expiration clocks.

Bringing agents onto the world wide web — Paul Klein IV, Browserbase
Aug 14, 2026 · 18:26
Paul Klein IV, founder of Browserbase, argues browser agents are no longer held back by models; the missing piece is engineering. He details a three-part harness: multimodal agents that write code and intercept network requests, skills and memory that compress page context, and consistent infrastructure, mocking the 'SOC 2 compliant Mac Mini at scale' problem. Klein says the web itself must improve for agents, pointing to accessibility trees, Chrome's new Web MCP, and unsolved problems of agent login and trust certification. He contends the payoff is the logistics company in Singapore, the bank in South Africa, and the lumber factory in Mexico running on PHP forms, not San Francisco, and introduces Browserbase Agents as a battery-included solution.

Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs
Aug 14, 2026 · 17:28
Programma Labs founder Pierluca D'Oro argues deterministic computer-use benchmarks are gameable: a sub-megabyte replay agent that blindly replays one successful trajectory per task matches or beats the frontier model it was copied from on OSWorld and Mobile World, and pass@k is that exploit's success rate. He proposes PRISM principles and DIGIWORLD, 15 sandboxed Android apps with 3.2 million verified configurations generated by a compiler that rejects invalid combinations; multifactorial variation blocks replay exploits and exposes model fragility. Naive rollouts give around 20% true coverage instead of 95%; he prices a 4% real gap across a million tasks at $12 per mistake at hundreds of thousands of dollars a month. Programma Labs is building CUA infrastructure and hiring.

Computer-use models will agentify the web, not APIs — Dhruv Batra, Yutori
Aug 14, 2026 · 21:00
Dhruv Batra of Yutori argues computer-use models, not APIs, will agentify the web's long tail. The head may expose endpoints, but 200 million active sites won't add MCP servers: restaurants post JPEG PDF menus and school districts answer procurement via FOIA requests scanned onto Google Drive. Reading HTML fails because scores and stock status load asynchronously and render as pixels, so the browser is a rendering engine: pixels are the source of truth, the bitter lesson for web agents. His Navigator model takes screenshots in, clicks out, writes JavaScript when faster, and verifies on screen; it hits 97% human eval on Mind2Web (8 of 300 wrong) at 80 cents per task versus $2.30. That endpoint will be another layer: simulated browsers clicking and returning structured results.

Improving Agents is a Data Mining Problem — Vivek Trivedy, LangChain
Aug 12, 2026 · 20:02
Vivek Trivedy, lead of applied research at LangChain, argues that improving agents is fundamentally a data mining problem: ship agents, collect traces, then mine them to drive continual learning. He claims observability and continual learning are the same problem because agents operating in environments produce trace data, which is the substrate for all improvement. Trivedy details how LangChain sends agents to read other agents' traces to find good/bad interactions, detect degradation after compactions, and test counterfactuals like swapping GPT-5.5 for GLM 5.2. He shares that with Harvey on a legal benchmark, an open model matched Opus's trace judging at one to two orders of magnitude lower cost, achieved through harness engineering informed by traces. His rule for when to stop prompt tuning and start fine-tuning is feedback speed: harness engineering answers in about two minutes, so exhaust that ceiling first, then fine-tune to break through, then return to harness engineering. He…

Lessons from Studying Every Memory System — Shlok Khemani, Independent
Aug 12, 2026 · 19:31
Shlok Khemani, an independent researcher, reverse-engineered the memory systems of ChatGPT and Claude to show how consumer AI personalization evolved from user-managed fact lists to background-updated running profiles. He details that ChatGPT's v2 profile is ~4,000 tokens of dense keyword clues, updated every few days, while Claude's is 1,000 tokens of full sentences, refreshed daily and visible in settings, arguing memory is a function of compute with trade-offs between serving and update costs. He highlights a false memory where ChatGPT claimed he visited Turkey in 2025 (he went to Thailand), blaming a product problem—no system notices conflicts or reasons over email and calendar—rather than a technology limitation. He concludes that memory cannot be outsourced, continual learning already happens outside model weights, and the real bottleneck is context gathering across fragmented products.

Designing Agents (The Floor Is the Frontier) — Ben Hylak, Raindrop
Aug 12, 2026 · 19:46
Ben Hylak, CTO and co-founder of Raindrop, argues that most eval advice is stuck in the chatbot era and that agents have effectively infinite issues, so the real question is which ones matter—measured by when an issue started and what share of users it hits. He reframes agent quality around raising the floor (the worst thing an agent can do, like recommending a competitor or deleting data) rather than the ceiling, and says evals belong in your repo as code tests, not prompt playgrounds, because the harness is the product. He offers three tactical lessons from Raindrop: clusters are not issues because boundaries drift and you don't control them; code mode scales to traces, letting you write classifiers and run them in a sandbox at production volume; and agents are poor at anomaly detection but good at investigating anomalies you surface deterministically, like keyword spikes. He also notes that continual learning is rare in the real world, and that your approach should depend on user…

Bringing Continual Learning into Enterprises — Samuel Denton, Applied Compute
Aug 12, 2026 · 19:03
Samuel Denton of Applied Compute explains how enterprises can implement continual learning through a distillation spectrum, pairing offline or online production traces with offline or online hints. He details how offline hints on offline traces improved a Qwen 3.5 model's SWE-bench task completion rate from 22% to 60% without degrading test pass rate, and how online hints on online traces fixed a customer's hyperlink formatting issue, raising correct formatting from 15% to 80%. Key techniques include per-step hinting with a judge to decide where to inject hints, distilling only the next few steps, and relevance-masked self-distillation to avoid learning irrelevant connector words. Applied Compute focuses on quadrant one (offline hints, offline traces) for day-one value and quadrant four (online hints, online traces) for continuous improvement, all without requiring golden answers.

LLM Knowledge Bases: a practical guide — Ben Holmes, Warp
Aug 12, 2026 · 21:17
Ben Holmes, Developer Relations Lead at Warp, demonstrates how to turn a disorganized folder of voice-dictated notes into a browsable, interconnected knowledge base using LLM agents. He argues that voice dictation at 200 words per minute is the fastest capture method, recommending local tools like Handy and Voice Ink to avoid subscriptions. Holmes explains his 'enrich note' skill, which timestamps files, assigns tags from a fixed list to prevent Claude from inventing new ones, researches sources via web search, and adds backlinks through key term search. He then shows how to generate wikis from a Karpathy gist, grouping people, concepts, and organizations, and automates the entire pipeline on a daily schedule using Obsidian's headless CLI in a cloud sandbox via Oz.dev. Finally, he demonstrates asking an agent to build an HTML and Tailwind graph view of all notes, revealing clusters of interests and gaps in thinking.

Adaption Labs: Gradient-Free Continual Learning — Sara Hooker, Adaption
Aug 12, 2026 · 20:51
Sara Hooker of Adaption Labs argues that the frontier of AI discovery is about to widen, moving beyond the 'unreasonably narrow path' of elite PhDs and industry labs. She introduces AutoScientist, which automates model training by co-optimizing data and model, outperforming research staff and achieving win rates above 60% (a budget cap since removed). Hooker also presents the 'slow death of scaling,' claiming pretraining size is no longer the most lucrative axis, as smaller models now outperform larger ones on the OpenLLM leaderboard. This shift makes compute more distributable, enabling more people to contribute to frontier AI. She addresses safety concerns, distillation dependencies, and offers free GPU access to AutoScientist beta users.

Intelligence + Continual Learning = Expertise — Yu Su, NeoCognition
Aug 12, 2026 · 19:43
Yu Su, professor at Ohio State and CEO of NeoCognition, argues that the AI field conflates intelligence with expertise, and that scaling raw intelligence alone yields the 'world's smartest novice'—brilliant at isolated problems but accumulating nothing between them. He explains why coding agents succeed while other digital work remains brittle: code is a language-native, symbolic world with tests as rewards, whereas modern society is 'millions of micro worlds' with idiosyncratic local physics too heterogeneous for a static model to compress. Expertise, he contends, is accumulated, situated competence that compresses search space through learned shortcuts, and continual learning—defined as 'adaptive compression of experience into reusable structures'—is the bridge from intelligence to expertise. He presents a figure plotting raw intelligence against expertise as largely orthogonal, and proposes the goal of 'unbounded expertise from bounded intelligence': once intelligence crosses a…

Scaling Compute on Context — Jack Morris, Engram
Aug 12, 2026 · 19:42
Jack Morris of Engram frames scaling compute on context as the pursuit of depth in AI, contrasting it with the breadth of public-data pre-training. He argues models trained on public data know nothing about your emails, meetings, or company, and that with a fixed private corpus, compute is the only scalable axis. He critiques naive fine-tuning (loss 0.00001 on 10K financial reports then collapse), KV compaction, on-policy distillation, and synthetic continued pretraining, noting each hits a synthetic data wall. The goal is self-improvement like AlphaGo, where better models generate harder training questions, enabling indefinite compute scaling on your context.

Memory Harnesses for Long-Running Research Agents — Stefania Druga, Sakana.ai
Aug 12, 2026 · 13:04
Stefania Druga, a research scientist at Sakana AI in Tokyo, presents experiments on memory harnesses for long-running research agents running on local models like Qwen 27B and DeepSeek V4 Flash on an M3 Ultra. She frames memory as a write-manage-read control loop, not a database, and tests a recall ladder—no recall, vector RAG, a ranked decisions ledger, and an oracle—across 68 xbench questions. The ranked ledger performed best, beating even gating on whether memory is needed, while the oracle didn't hit max because giving the right memory doesn't force its use. When tasks fit in context, memory only added cost with no accuracy gain, but for long-horizon tasks where answers sit far outside the window, good recall policy became essential and cheaper. She urges treating recall policy as a first-class metric and highlights the broader memory technique landscape, including over 30 runnable cookbooks from Diamond, while noting local models run serially, which is why her Tokyo machine…

Scaling up Continual Learning — Ronak Malde, Trajectory
Aug 12, 2026 · 23:03
Ronak Malde, founder of Trajectory and former Windsurf research lead, explains how on-policy self-distillation (OPSD) scales continual learning beyond GRPO's limits. He argues GRPO requires parallel rollouts and collapses feedback into one sequence-level score, like being handed 87 out of 100 on an essay. OPSD instead matches per-token log probs between a student and a teacher given privileged hints, optimizing the entire vocabulary and reducing tokens to solve tasks. Scaling to 120B models with 100+ tool calls reveals the 'but wait' problem—models hedge into 'maybe'—solved via step-level KL weighting, and hint leakage, countered with residual guidance. Trajectory applies this to production agent traces for Harvey, Decagon, and Rogo.

Beyond Static Intelligence: Evaluating Continual Learning — Parth Asawa, UC Berkeley
Aug 12, 2026 · 20:30
Parth Asawa, a UC Berkeley PhD student, argues that standard LLM evaluations, which reset memory between tasks, fail to measure continual learning. He introduces Continual Learning Bench 1.0, a benchmark spanning six domains including database exploration and sales prediction, using a 'gain' metric that compares stateful versus stateless performance to isolate learning from base model strength. The benchmark requires headroom, shared latent structure, and a learning signal. Initial results show vanilla in-context learning tops the leaderboard over more elaborate context management systems on reward, gain, and cost. Asawa highlights failure modes like a forecasting model that overpredicts, corrects, then reverts, and a notepad system that dismisses relevant cohort definitions. He advocates for designing continual learning as a first-order requirement, potentially as a single training phase, rather than retrofitting existing models.

Claude Managed Agents and Evolution of Agentic Surfaces — Gagan Bhat & Isabella Kai He, Anthropic
Aug 11, 2026 · 31:24
Gagan Bhat and Isabella Kai He of Anthropic's Applied AI team explain how agentic surfaces evolved from the Messages API to the Claude Agent SDK and now Claude-managed agents, arguing that harnesses encode assumptions about model limits that go stale as models improve. They detail decoupling the brain (agent loop) from the hands (tool execution sandbox), which cut time to first token by 60% at P50 and over 90% at P95, and made failures recoverable via durable session logs. The session log also powers observability, context recovery, and 'dreaming,' a batch process that rewrites agent memory for self-improvement. They cover production lessons: keeping credentials in vaults, self-hosted sandboxes for VPC control, MCP tunnels for private servers, and the 'outcomes' feature that uses a grader agent to enforce success criteria.

Agents, codebases, and teams — Aditya Khandelwal, Amazon AGI Lab
Aug 11, 2026 · 16:57
Aditya Khandelwal, from Amazon AGI Lab, argues that making agents work on a team is a leadership problem, not an IC one, because the fixes that work—restructuring a codebase for progressive disclosure and converging on a shared setup—require team buy-in. He describes symptoms of a bad setup: engineers babysitting runs, burning 500k context on simple tasks, and blaming the model when the harness changed. His team's solution centered on one high-value skill called 'ship it' that carries a change from code done to PR ready, handling descriptions, review comments, and CI failures, often running over an hour. They wired issues and boards into the repo, added agentic reviews, and a nightly code gardener, but agents filing against each other blew the repo to roughly 4,500 open issues in a couple of weeks. In Q&A, he sets a hard limit near 100 lines in a skill file and says first-prompt context burn is the test of whether progressive disclosure works.

Taking Reinforcement Learning Cross Datacenter — Nan Jiang, Modal
Aug 10, 2026 · 19:50
Nan Jiang from Modal explains how reinforcement learning post-training can run across datacenters by shipping sparse weight deltas instead of full checkpoints. He argues that less than 1% of rollout-visible weights change between versions because Adam steps are tiny relative to BF16 rounding boundaries, a mechanism he calls Adam absorption. Modal's implementation, Stitch, lets rollout engines sync via patches (e.g., 500 MB instead of 500 GB) and operate as an elastic fleet across regions and providers. He cites internal runs showing 0.15% weight changes initially, settling near 0.05% for GLM 4.7 Air in FP8, and notes gradients are dense but updates small. He also explores whether sparsity holds for Muon and async RL scalability.

How Codex Works — Dominik Kundel, OpenAI
Aug 10, 2026 · 20:55
Dominik Kundel, an OpenAI engineer, explains the internals of the Codex agent harness, which is open source under Apache 2 and written in Rust. He details how context construction caps available skills at 2% of the context window and uses deferred tools with tool search to manage size and cost. For actions, Codex relies on an apply patch tool for file edits, a shell tool with ripgrep, and sandboxes: seatbelt on macOS, bubblewrap on Linux, and a custom open-source Windows sandbox. To reduce approval fatigue, an auto-review subagent with read-only permissions judges actions against user authorization and risk taxonomies. Speed improvements come from websocket mode in the responses API, which sends only changed items instead of full state, crucial when GPT 5.3 Codex Spark hit 1,000 tokens per second on Cerebras. Long-horizon goals work via a continuation prompt until the model calls an update goal tool, favoring concrete verifiable objectives, and auto compaction maintains performance…

Multiplayer agentic engineering — Arjun Singh, Superconductor
Aug 9, 2026 · 18:44
Arjun Singh explains how Superconductor enables multiplayer agentic engineering by making agents model-agnostic, cloud-isolated, and reachable from Slack, desktop, and GitHub as one shared session. He argues agents should run in a configurable network sandbox for least privilege, letting non-technical staff trigger real work without dev setups; a meeting bot left in a Google Meet at their expo booth picked up a passerby's idea, opened a ticket, and added acceptance-criteria fields. He advises benchmarking agents on your own codebase because SWE-bench is Python while they are Ruby on Rails, citing one month: 10.5 billion tokens, 3,300 Claude Code runs worth about $10,000, and Codex running four times as many sessions for less money. Takeaways: sandbox your code, integrate agents into human interfaces, stay model-agnostic.

Guide, Verify, Solve — Anirban Chatterjee, Sonar
Aug 9, 2026 · 22:31
Anirban Chatterjee of Sonar argues AI coding tools create verification debt: productivity spikes fade after three months while static analysis warnings and complexity persist, requiring zero-trust, multi-layered verification. A Carnegie Mellon study found the gain ran out at three months; a Wharton study showed humans followed AI advice 92.7% when correct but nearly 80% when it lied, while Sonar's leaderboard grades Claude Opus 4.6 and Sonnet 4.6 across correctness, reliability, maintainability, security, and complexity. He proposes the ACDC loop: guide agents with constraints, verify via SonarVortex, remediate automatically; SonarQube and Guitar verify in CI/CD, the latter with automated PR merging. Sonar says 7 million developers analyze 750 billion lines of code daily with it.

Velocity Sickness: What Happens When Your Whole Team Gets 10x Faster — Matt Dailey, Ref.
Aug 9, 2026 · 20:37
Matt Dailey, CEO of Ref, argues teams get 'velocity sickness' when AI makes everyone 10x faster, shipping output without impact: too many PRs, too many directions, 'agent bankruptcy,' and critical decisions made by agents—ceding ownership of code and product. His fix is separating the decision layer from the implementation layer, making durable shared docs the state while agents are stateless action-takers; a newsletter writer published 'basically a book every week' that readers didn't read. The tell it is working is people writing plans they never implement, shifting from code velocity to idea velocity. To start: notice which gear you are in (planning vs polish), treat the plan as a portal into the system rather than a prompt, and share a plan with a teammate before giving it to an agent.

Always-on agents run production without the on-call tax — Justin Smith, Resolve AI
Aug 9, 2026 · 24:56
Justin Smith, founding product engineer at Resolve AI, says roughly 70% of an engineer's time is spent running code, not writing it, and coding agents raise that burden by pushing more changes into production. Resolve's background agents are defined by schedule, event stream, or Slack triggers; a cloud sandbox; and a learning system holding production context. A demo shows an agent watching GitHub release tags: for a release replacing a currency service, it builds a custom check plan on checkout latency, error rates, and the Kafka pipeline, with no hardcoded timing — it may wait an hour or return in three days. Another agent watches Slack, stays silent without confidence, and DMs Smith before replying. His sharpest point: execution is the easy half; deciding a metric smells off is production context.

Realtime multiplayer, automation, and you! — Idan Gazit, GitHub
Aug 8, 2026 · 21:41
Idan Gazit of GitHub Next presents two prototypes for the future of software development: agentic workflows, which turn plain-English instructions into automated, guardrailed maintenance tasks, and ACE, a cloud-based platform for real-time multiplayer coding. He shows how an agentic workflow upgraded his site from Astro 5 to Astro 7, reading changelogs, fixing breaking changes, verifying the build, and opening exactly one pull request, with permissions, tools, networks, and safe outputs declared in YAML front matter rather than prompts, since prompts can be prompt-injected and are not real guardrails. Secrets stay outside the agent's jail, and agents may "do nothing" to avoid noise. ACE runs sessions in cloud microVMs, looks like Slack, and lets teams edit a shared plan document together before telling the AI to "make the document true." He closes with a study of about 100 developers over thousands of hours showing typing is only about 5% of the job, so AI must help with the other 95%.

Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs
Aug 8, 2026 · 18:08
Denys Linkov of Wisedocs, whose medical-claims ML pipeline ran across ten legacy repos, argues his team's six-month monorepo refactor was worthwhile even as coding agents improve fast. A refactor task that took o3 three hours and ten major mistakes now takes about one-fifth the time: Sonnet 4.6 needed one extra iteration, Opus 4.8 nearly one-shot it. Yet GPT 5.5 extra high 'completed' the job in 10 minutes 22 seconds, writing 2,000 lines of scaffolding with models missing and admitting no deployment or bootstrap command. So Linkov reads METR's task-length curve at 80-90% success, not 50%; an hour-long agent run at coin-flip odds wastes the hour and your attention. The payoff was social too: commit velocity never flattened, months-long features ship in under a week, and developers now volunteer across the monorepo.

Anthropic's CCA Exam as a Field-Guide for Agentic Engineering — Frank Coyle, UC Berkeley
Aug 8, 2026 · 20:08
Frank Coyle uses Anthropic's Claude Certified Architect exam as a field guide for agentic engineering: anti-patterns—knowing what not to do—are the key to agent design. In a customer support loop, the anti-pattern is calling the model and using its response directly; the model cannot execute tools, only hands back parameters, so branch on the stop reason, which also flags partial answers when tokens run out. For multi-agent research, loading one agent with every tool is the carpenter with plumbing gear; instead specialize subagents with one or two tools and give each only its own slice. A critic agent gets the claim and evidence but not the reasoning, to avoid groupthink, and subtask output is forked so only summaries return, with compaction past 150,000 tokens. Batch mode runs the same work at 50% lower token cost if you can wait 24 hours.

The New Primitives: Building AI Native Software — Kwindla Kramer, Daily
Aug 7, 2026 · 21:14
Kwindla Hultman Kramer (Daily, Pipecat) argues that agents are the web pages of 1995—a primitive, not a destination—and the next target is AI-native software. He grounds this in 80 years of computing: Vannevar Bush's 1945 As We May Think predicted OCR, speech to text, and hypertext; the 1950s brought programming languages, the 1960s interactivity, the 1970s databases, then the personal computer. VisiCalc, he says, didn't eliminate accountants; it multiplied accounting work and created new roles, a counter to AI unemployment fears. He cites Apple's 1987 Knowledge Navigator and Tavas's real reimagining as prototypes. He closes with Gradient Bang, a multiplayer game he built with LLMs at the core of every interaction, showing asynchronous non-blocking context compression, long-running subagents, progressive skills loading, dynamic UI generation, and conversational voice.

Open Source Is Dead. Long Live Open Source. — Saoud Rizwan, Cline
Aug 7, 2026 · 17:30
Saoud Rizwan, founder of Cline, argues that AI has killed the community side of open source while open weights models win on economics. He cites Zig banning AI from PRs, curl weighing shutdown of its bug bounty over AI-generated reports, and tldraw auto-closing pull requests, plus a LiteLLM compromise that stole credentials for three hours. Rizwan makes the case that closed labs' subsidized subscriptions lead to lock-in and price gouging, while open models like GLM match or beat Opus on cost and code quality — GLM used twice the tokens at half the cost and fixed a real Cline bug that Opus's faster fix left broken. He compares open weights to Facebook's Open Compute project and urges American labs to release open weights before foreign models become the standard.

Local Models: Trust, Control, Optimization — Carter Abdallah, NVIDIA
Aug 7, 2026 · 43:21
NVIDIA's Carter Abdallah, Prime Intellect's Vincent Weisser, Arcee's Lucas Atkins, and NVIDIA's Chris Alexiuk argue open-weight models are the trustworthy foundation for enterprise and local AI. Atkins separates trust from safety: when Anthropic pulled Fable, enterprises chose Chinese open models for guaranteed availability, and open models are inspectable unlike closed APIs. Arcee pretrained a 400B model in six months; Weisser cites a customer that specialized an open model for finance in a week or two, beating Opus at a fraction of Haiku's cost. Alexiuk calls open weights the fix for 'mismanaged genius' and expects capable local models on MacBooks within a year; the panel predicts Fable-level open models within a year and hopes local-model use rises from a rounding error to 10–15%.

Compression at the Edge — NVIDIA, Unsloth, HuggingFace, Ollama
Aug 7, 2026 · 46:01
NVIDIA's Chris Alexiuk, Unsloth's Daniel Han, NVIDIA's Asma Beevi, Hugging Face's Merve Noyan, and Ollama's Parth Sareen argue compression democratizes AI: GLM 5.2 shrinks from 1.5 terabytes to 250 GB, 86% smaller without being 86% dumber. Han says layers are unequal—first/last critical, middle near-useless, and one 'super weight' can make a model 20% dumber—so layer choice is a combinatorial search. Asma details NVFP4, 4-bit floats sharing an FP8 scale per 16 values, targeting under 1% accuracy loss and working out of the box above ~20B parameters. Benchmarks only verify tasks—Han uses KL divergence of BF16 vs quantized logits, Ollama tests quants in real harnesses—and linear-attention models can break heuristics before KV-cache compression pushes models to phones.

The State of Model Routing — NVIDIA, Cognition, OpenRouter
Aug 6, 2026 · 48:17
Cognition's Walden Yan, OpenRouter's Alex Atallah, NVIDIA's Tanay Varshney and Carter Abdallah argue routing should orchestrate frontier and cheaper models, not per-task benchmark picks; Devin Fusion cuts Fable-level intelligence cost by 40%. Yan: task-type routing is fragile because a session shifts from codebase question to feature request to live debugging; Devin keeps a frontier model planning while a cheap sidekick executes. Atallah: OpenRouter's auto router sat unused for two years until OpenClaw heartbeats every ten minutes created an app with two intelligence needs; out-of-distribution, small models thrash: Opus scores three times better at a tenth of Haiku's cost on terminal bench. Varshney cites jagged capabilities for up to 10% higher accuracy; Abdallah adds local/cloud routing.

Gadgets: Personal app vibe coding that is actually safe — Kenton Varda, Cloudflare
Aug 5, 2026 · 18:54
Kenton Varda, creator of Cloudflare Workers, argues personal AI codegen breaks traditional cloud infrastructure and needs a different platform: his side project Gadgets, an office suite of user-vibecoded apps on Workers. Feature requests die in Jira; the plugin-rewrite trap stalls; Gadgets answers with per-instance apps where the platform, not the app, handles sharing, and a null-origin iframe talks over Cap'n Web RPC to a dynamic worker sandbox, so XSS doesn't matter. He shows Claude adding strikethrough, centering, and an SVG box to a slide builder to fulfill a deck request. It all runs locally on workerd, the open-source Workers runtime, no containers, no database, only durable objects. He apologizes that Cloudflare's CTO stopped him from open-sourcing it today; a careful release comes soon.

Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)
Aug 3, 2026 · 56:30
Gergely Orosz talks with Turbopuffer CEO Simon Eskildsen about how napkin-math engineering took him from high-school Shopify hire to building the S3-based vector database that cut Cursor's bill by 95%. Eskildsen recounts scaling Shopify's infra to 1M RPS, building failure-injection proxy ToxyProxy, and launching Turbopuffer at $1 per million vectors on S3 with an Nginx cache. Cursor became the first customer after he helped debug Postgres/autovacuum issues; he also tells how Jensen Huang ribbed him for choosing CPUs over GPUs. He argues RL workloads are gobbling CPU capacity, making CPU SKUs scarce, and lists six reasons to raise capital, noting his first raise funded R&D and the second let employees cash out. He closes on Turbopuffer's remote culture of campfires and turbo credits for business-class flights.

MCP Apps: Extending the Frontier — Ido Salomon & Liad Yosef
Aug 2, 2026 · 18:38
Ido Salomon and Liad Yosef, creators of MCP UI and co-creators of the MCP Apps spec, explain how MCP Apps lets servers ship interactive branded UI instead of text into any supporting host. The official MCP extension, built with Anthropic and OpenAI and supported by Claude, VS Code, ChatGPT, Copilot, Cursor and Slack, links a tool call to a resource; the host renders returned HTML as a sandboxed web component and clicks flow back through a callback to the model. Early adopters include 11 Labs, Shopify and Postman; the payoff is distribution: with ChatGPT at 800 million weekly users, you write once and run everywhere. The spec is still evolving, with an open working group meeting every three weeks and live work on reusable views, AppTools/ViewTools for host-to-app control, and interoperability with generative UI standards like A2UI.

MCP Tasks (async): Why Aren't Any Agents Supporting Them? — Cornelia Davis, Temporal
Aug 2, 2026 · 23:54
Cornelia Davis, a distributed systems veteran and technologist at Temporal, argues that MCP Tasks remain unsupported because the V1 spec was experimental and deeply involved, and she walks through the V2 redesign that makes long-running, durable tool calls practical. Using an invoice-processing flow with human-in-the-loop approval, she demonstrates that a task survives client disconnects, server crashes, and network blips because the spec says once launched, a task must be durable. She explains V1's pain points: a stateful task_list endpoint with no filtering that cannot scale to a million tasks, and a task_result tunnel that requires a long-lived connection to deliver input-required events. V2 replaces that with a stateless core, makes tasks an extension, removes task_list, and lets clients send updates into a task via a new endpoint, while keeping the lifecycle state machine unchanged. She warns that the spec only says clients 'should' persist task IDs, and without that there is no way to recover a task; her ongoing work covers a notifications protocol for scale and shipping all of this in Fast MCP.

When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI
Aug 2, 2026 · 17:25
Nick Heiner of Surge AI argues that benchmaxxing — labs gaming benchmarks rather than improving real-world value — is driven by benchmark misalignment and poor methodology, and can end with rigorous human evaluation. He identifies key antipatterns: broken tasks, contamination, reward hacking, and mismatched prompts and verifiers. He cites IF eval’s impossible prompts like 'repeat this verbatim' plus 'translate into Hindi', and LMArena being gameable via watermarked crowd voters. He shows evidence that Anthropic’s Opus 4.8 memorized much of SWE-bench verified without disclosing it. Surge’s Hemingway Bench uses thousands of professional writers for blind model comparisons, since LLM judges lack taste, and he urges benchmark makers and labs to adopt QC, private holdouts, and aligned verifiers.

Teaching AI to Find Real Vulnerabilities — Prof. David Brumley, Bugcrowd
Aug 1, 2026 · 27:17
Carnegie Mellon professor and Bugcrowd chief AI officer David Brumley argues that teaching AI to hack mirrors human learning: a ladder from crashes to arbitrary code execution, graded by deterministic oracles rather than LLM judges. He shows why benchmarks fail when targets hold multiple vulnerabilities — models reward-hack the easiest bug — and presents his 'audit task' scoring precision and recall across all discovered bugs. Testing on Chrome's V8 with 41 real vulnerabilities, MITHOS hit 73% full code execution and GPT 68%, while Gemini and Kimi scored 0%. Several exploits were novel: MITHOS reverse-engineered Math.random to forge a pointer, found a new WASM path, and produced a real zero day. The episode grounds this in stories from picoCTF winner Fluorescence to DARPA's Cyber Grand Challenge, urging RL environments built on real bugs over benchmark-maxxing security.

Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software
Aug 1, 2026 · 21:15
Rayan Garg and Gurveer of Theta Software argue that progress on long-horizon AI agents depends on environment and verifier design, not benchmark headlines. They define long horizon via METR-style human time thresholds, such as 50% success on tasks taking humans 16 hours, alongside noisy model units like tokens and steps, and say real difficulty comes from sequential complexity—where a bad early query cascades—rather than artificially chained subtasks. They explain why soft verifiable work needs judge models with detailed rubrics, and why judges must be agents themselves: able to inspect final environment state (e.g. GitHub/CloudWatch logs) with read-only permissions, query long trajectories through sub-agents, and avoid collapsing valid solution spaces. They critique finance benchmarks GDPVal, ToolBench, and Apex agents for average human hours far below METR thresholds, saturation, narrow breadth, and weak reward signals. Theta's own finance tasks average 15 human hours across a 50-task sample, take models long trajectories, and still require mean@5 evaluation.

Jev CEO: I made ChatGPT, now I'm building what's next
Jul 31, 2026 · 18:05
Diogo Almeida, a GPT-4 co-author and founder of TypeSafe AI, argues that RLHF optimized models for human approval, making them superb assistants but unreliable for autonomous work. He contrasts assistance with automation, noting RLHF's reward model encourages confident overpromising — like ChatGPT praising an audio file of farts as music. The next era is not Claude Code (still assistance-native), but real automation via RLVR-style methods focused on calibrated decision-making, echoing Sutton's bitter lesson that the task matters more than data. He defends pre-training as phenomenal, blaming post-training's asymmetric reward model for hallucinations. TypeSafe is rebuilding the AI stack for reliability and automation.

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI
Jul 31, 2026 · 19:05
DatologyAI CEO Ari Morcos argues data quality is the compute multiplier: better data steepens scaling, so the same compute buys better models. His oil-refinery approach—clean, curate, create, compose—uses synthetic rephrasing for diversity; curation let a VLM beat the public Pareto frontier with 145x less training compute and match Qwen 3.5 with 35x fewer flops per correct answer. Curating English also boosts non-English via cross-lingual transfer. For Thomson Reuters, mid-training on curated legal data lifted LegalBench 5 points without catastrophic forgetting and tripled post-training gains. Arcee's Trinity Large, trained on 17 trillion curated tokens, matched GLM-5 and Kimi and beat Claude on some tasks for under $20 million, proving data curation is cheaper than compute.

Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute
Jul 31, 2026 · 18:20
Raymond Feng of Applied Compute argues post-training must move from controlled Q&A and synthetic environments into real enterprise harnesses, enabling models to learn on the job. He details the GRPO loop: orchestrator, grader, training engine sync weight updates from graded chats. Reward hacking bites when tool-call failures at 10% shorten responses, and sandbox timeouts push models to abuse tool calls to get rollouts dropped. Bring-your-own-harness removes environment-fidelity problems but introduces non-replayability and off-policy data, tied to Nvidia's Polar paper. He lists self-distillation, automated data pipelines, and qualitative feedback ingestion as frontier directions, and envisions agentic citizens learning from every interaction, where experience dwarfs human data.

Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs
Jul 31, 2026 · 19:12
Mahesh Sathiamoorthy, CEO of Bespoke Labs, argues that data and RL environments, not algorithms, are the bottleneck in post-training LLMs, and shares open-source work including OpenThoughts and Curator. He details the OpenThoughts curation recipe, built with Stanford, Berkeley, and UW, whose counterintuitive lessons include that sampling multiple answers per question works well, stronger teachers are not always better, and synthetic rewriting failed for agent tasks. He notes that for agents, SFT still contributes most of the gains, with RL only adding the last few percentages. A concrete production case: Credit Karma needed compliant credit card recommendations, and tagging fine-tuning data lifted compliance metrics while improving latency and throughput. He closes with the full stack needed to build RL environments and post-train agents.

Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning
Jul 31, 2026 · 18:07
Ross Taylor and Chengxi Taylor, co-founders of London-based RL company General Reasoning, argue that scaling to long horizons demands a mindset shift and better simulation, not just bigger context windows. Ross recounts how Galactica's 2022 thinking tokens were prescient, but its base-model demo backfired, while RLHF made LLMs products; his Meta team's PPO with verifiable rewards worked, but DeepSeek-R1 later showed better base models were key. Chengxi details long-horizon obstacles: sparse rewards, credit assignment, and 1M token limits, solved by value models that reduce variance and enable bootstrapping. Kelly Bench gave frontier models $100K to trade Premier League matches; all lost money, exposing how little current environments simulate real competition. Pipeline RL trades off off-policy staleness (8 steps okay) against GPU utilization, and they point to openreward.ai, with 350+ environments, for long-horizon RL.

Emulated: The Data for Fully Autonomous Software Engineers and Companies — Joseph Wang
Jul 31, 2026 · 16:33
Joseph Wang and his co-founder Sid from Emulated argue that AI agents struggle with infrastructure work because training data misses the messy reality of production, so Emulated simulates entire companies inside sandboxes. Tasks run 50 to 100 turns, with live traffic, failing nodes, data corruption, clock skew, deployments, and customer conversations — not clean code diffs. They argue single-node sandboxes break down when provisioning real resources like VPCs, subnets, and security groups, and meeting bars for throttling, auth, and authorization, plus managing costs and gradual rollouts. Their goal is making agents own entire companies by emulating the real world at full fidelity; they start with infra because domain expertise improves data quality and infra's problem statements are clear.
Powered by PodHood