A company discussed on AI Engineer.

Can LLMs Write Fast Multi-GPU Kernels? — Simran Arora, Together AI
Aug 27, 2026 · 30:00
Simran Arora, principal scientist at Together AI, argues multi-GPU communication, not single-GPU compute, is AI's bottleneck, and introduces ParallelKittens and ParallelKernelBench to test whether LLMs can write fast multi-GPU kernels. From A100 to B200, BF16 tensor core throughput improved 7.2x vs 3x intra-node communication, leaving PyTorch+NCCL baselines below 50% of communication-aware roofline on most problems. ParallelKittens adds a dozen lines to a single-GPU kernel and runs in production at Together and Cursor; in ParallelKernelBench's 87 problems, the best model solves 28 zero-shot (22 faster), and more samples lift correctness to 36 but fast-and-correct stalls near 31%. Models compile after retries but stall on ordering and transfer choices; a bash-based agent hits 35 correct.

Einstein Arena: Harnessing Collective Agent Intelligence for Open Science — James Zou, Together AI
Aug 25, 2026 · 16:56
James Zou, Stanford professor and Together AI collaborator, argues that designing environments, not workflows, unlocks collective AI agent intelligence, presenting Einstein Arena and DSGym. In Einstein Arena, agents prove they are bots and collaborate on open scientific problems; within weeks they found best-known answers to 11 problems, including raising the 11-dimensional kissing number from 593 to 604. The same arena, with kernel compilation as verification, produced kernels over 2x faster than the prior state of the art, now in production at Together AI. DSGym found that 20-50% of tasks in popular benchmarks could be solved without touching the data; frontier models still score under 50%, while execution-verified trajectories fine-tune open-source models that run on a laptop.

The Missing Layer: Design Taste in AI Agents — Hassan El Mghari, Together AI
Aug 21, 2026 · 14:10
Hassan El Mghari, developer experience lead at Together AI, argues design taste, not model capability, decides whether AI apps reach millions or read as vibe-coded slop. He ships about 10 apps a year — a few with millions of users — and credits design for that reach. His design skill Hallmark, launched six weeks ago to 10,000+ users, turns AI slop tells like purple gradients and italic headers into gates and supplies themes, since strong inspiration produces better work. He starts in Codex or Claude Code and iterates with cheaper open-source GLM 5.2, whose landing page was nearly indistinguishable from Opus 4.8. He advises saving screenshots as references, using longer prompts from voice notes, one or two features per prompt, and a skill file or AGENTS.md — but agent output is only a base.

Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It — Dan Fu and Olive Song
Jul 31, 2026 · 20:14
Olive Song, RL lead at MiniMax, and Dan, Together AI's VP of Kernels, argue that open-weight MiniMax M3 is closing the gap to frontier labs via community feedback and inference optimization. Dan details the day-zero stack: GPU kernels for M3's sparse attention, KV cache at million-token contexts, and adapting to agentic workloads that upload codebases. Song says M3 trains multimodal from scratch to prevent text-vision collapse, and uses RL with environment design and rewards for long-horizon tasks like replicating a 12-hour iClear paper, plus validation/test splits to avoid hacking. Together's Parallel Kernel Bench holds unsolved problems, so overfitting it yields kernels the company deploys. Both see self-evolution and better GPU utilization as why open models keep accelerating.

Road to 5 Million Tokens: Breaking Barriers in Long Context Training — Max Ryabinin, Together AI
Jun 8, 2026 · 15:50
Max Ryabinin from Together AI presents their research on extending transformer context length to 5 million tokens using Untied Ulysses, which cuts activation memory by reusing buffers across attention head iterations. The talk walks through a stack of techniques including fully sharded data parallelism, DeepSpeed Ulysses context parallelism for an 8x activation reduction, activation checkpointing for another 8x, CPU offloading of transformer block inputs, and chunked sequence training. Even with these, training a LLaMA 3B model with 3 million tokens fits on an 8xH100 node, but 5 million requires Untied Ulysses. Instead of allocating one large buffer per attention head group, it chunks heads further and reuses buffers across iterations, cutting activation memory with negligible throughput impact. At both 8B and 32B scale, results match the most memory-optimized transformer training baselines while pushing sequence length 25% further than prior Ulysses implementations.

Engineering voice agents: Latency, quality, and scale — Rishabh Bhargava, Together AI
May 31, 2026 · 24:35
Rishabh Bhargava from Together AI outlines engineering voice agents, explaining that pipeline architecture with colocated models can achieve sub-500ms response times critical for user retention. Speech-to-text targets P90 under 100ms and 6% word error rate, while the LLM must stay within 200-300ms time-to-first-token using 8-30B parameter models—larger models blow the budget, smaller ones break tool calling. Network latency from distant data centers adds 75ms (30% overhead) versus 5ms when colocated in the same building. Pure speech-to-speech models are emerging but still struggle with instruction following and tool calling. The thinker-talker pattern uses a small LLM for fast conversational flow and issues a single tool call to a larger model for complex requests.

From Mixture of Experts to Mixture of Agents with Super Fast Inference - Daniel Kim & Daria Soboleva
Jun 27, 2025 · 53:15
Daria Soboleva and Daniel Kim of Cerebras explain how Mixture of Experts (MoE) architectures enable scaling large language models efficiently by replacing monolithic feedforward networks with specialized experts, a technique used by GPT-4 and Claude. They then introduce Mixture of Agents (MoA), which combines multiple LLMs with custom prompts to outperform frontier models like GPT-4o on complex tasks, reducing a 293-second reasoning problem to 7.4 seconds using Cerebras' ultra-fast inference. The workshop guides participants to build their own MoA system, configure agents for bug fixing and performance optimization on a Python function, and achieve scores up to 120/120. Daniel details Cerebras' wafer-scale chip with 900,000 cores and distributed memory that eliminates memory bandwidth bottlenecks, enabling linear scaling and 15.5x faster inference on Llama 3.3-70B versus GPUs. Daria discusses ongoing research in diffusion models and sparsity, while Daniel notes plans for multimodal APIs and LoRA fine-tuning support.

Navigating AI’s Frontier in 2025 - Grace Isford, Lux Capital
Mar 13, 2025 · 17:55
Grace Isford, partner at Lux Capital, argues that while 2025 is a 'perfect storm' for AI agents with reasoning models like O3 and R1, cheaper inference, and billions in infrastructure (e.g., Stargate, DeepSeek), agents still fail due to cumulative errors—decision, implementation, heuristic, and taste—exemplified by OpenAI Operator booking a flight incorrectly. She prescribes five strategies: curating proprietary and agent-generated data, building personalized evals for non-verifiable domains (e.g., seat preference), designing scaffolding that prevents cascading failures (citing Ramp's approach), treating UX as the moat (e.g., Codium, Harvey, TLDraw), and building multimodally with voice, smell via Osmo, and touch for embodiment. The talk, recorded at the AI Engineer Summit 2025 in NYC, closes with a call to reframe perfection through visionary product experiences.

Accelerating Mixture of Experts Training With Rail Optimized InfiniBand Networking in Crusoe Cloud
Feb 12, 2025 · 17:45
Ievgen Bakulenko, product manager at Crusoe Cloud, explains how their rail-optimized InfiniBand networking accelerates training for sparse mixture of experts models. By leveraging NVIDIA's PXN feature, which allows GPUs to communicate across different rails using the internal NVSwitch in a single hop, Crusoe achieves a 50% improvement in synthetic benchmark latency and bandwidth for both small and large messages. In a real-world test fine-tuning the Mixtral model (8 feed-forward blocks, 7 billion parameters) on 240 H100 GPUs, this topology reduced training time by 14%, directly lowering cost and time-to-train. Bakulenko also outlines Crusoe's AI cloud platform, its climate-aligned mission using stranded energy, and its focus on easy-to-use infrastructure for AI engineers.

The GenAI Maturity Curve or You Probably Don't Need Fine Tuning: Kyle Corbitt
Feb 9, 2025 · 18:03
Kyle Corbitt, CEO of OpenPipe, argues that most teams don't yet need fine-tuning and should start with prompted models like GPT-4. He presents a GenAI maturity curve where the trigger to fine-tune is when you hit constraints on cost, latency, or quality consistency—for example, if GPT-4 is 80-90% correct but inconsistent on the last 10-20%. Fine-tuning shifts the paradigm frontier outward, enabling models like fine-tuned LLaMA 38B to outperform GPT-4 at 1/25th the cost. The process has four steps: capture production logs to know your input distribution, prepare high-quality data (using GPT-4 outputs or iterative labeling), train with one-click tools, and evaluate with inner-loop (LLM-as-judge) and outer-loop (business metrics) evals. OpenPipe and other providers make deployment trivial via OpenAI-compatible APIs. The talk delivers a concrete decision framework and a walkthrough so any engineer can fine-tune in under an hour.
Powered by PodHood