A product discussed on AI Engineer.

Can LLMs Write Fast Multi-GPU Kernels? — Simran Arora, Together AI
Aug 27, 2026 · 30:00
Simran Arora, principal scientist at Together AI, argues multi-GPU communication, not single-GPU compute, is AI's bottleneck, and introduces ParallelKittens and ParallelKernelBench to test whether LLMs can write fast multi-GPU kernels. From A100 to B200, BF16 tensor core throughput improved 7.2x vs 3x intra-node communication, leaving PyTorch+NCCL baselines below 50% of communication-aware roofline on most problems. ParallelKittens adds a dozen lines to a single-GPU kernel and runs in production at Together and Cursor; in ParallelKernelBench's 87 problems, the best model solves 28 zero-shot (22 faster), and more samples lift correctness to 36 but fast-and-correct stalls near 31%. Models compile after retries but stall on ordering and transfer choices; a bash-based agent hits 35 correct.

Accelerating Mixture of Experts Training With Rail Optimized InfiniBand Networking in Crusoe Cloud
Feb 12, 2025 · 17:45
Ievgen Bakulenko, product manager at Crusoe Cloud, explains how their rail-optimized InfiniBand networking accelerates training for sparse mixture of experts models. By leveraging NVIDIA's PXN feature, which allows GPUs to communicate across different rails using the internal NVSwitch in a single hop, Crusoe achieves a 50% improvement in synthetic benchmark latency and bandwidth for both small and large messages. In a real-world test fine-tuning the Mixtral model (8 feed-forward blocks, 7 billion parameters) on 240 H100 GPUs, this topology reduced training time by 14%, directly lowering cost and time-to-train. Bakulenko also outlines Crusoe's AI cloud platform, its climate-aligned mission using stranded energy, and its focus on easy-to-use infrastructure for AI engineers.
Powered by PodHood