A product discussed on AI Engineer.

Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai
Aug 18, 2026 · 16:55
Gabriel Jorge Menezes of Krea explains how Krea 2 was trained from scratch on thousands of GPUs, arguing metrics, not mystery-solving, made scale survivable. He calls GPU utilization a lie, tracking tensor core utilization instead, which climbs as resolution steps from 128 to 1024. Any GPU above 78 degrees gets pulled without debugging; custom InfiniBand and NVLink error collection caught most failures, which were silent cross-node timeouts. Aggressive checkpointing on a Wekker filesystem writing nearly a terabyte per second let crashes restart and run 24 hours on same nodes. Production and training share one cluster: gang scheduling kicks inference to external providers via a Virtual Kubelet fake node; taints stop wasted GPUs; a descheduler migrates pods back so the site never drops.

The Small Model Infrastructure Nobody Built (So We Did) — Filip Makraduli, Superlinked
May 5, 2026 · 18:30
Filip Makraduli of Superlinked introduces SAI, an open-source inference engine for small models that addresses gaps in embedding infrastructure by enabling dynamic model loading, hot-swapping, and memory-aware eviction on a single GPU. He argues that provisioning separate GPUs for each small model wastes idle capacity, and that the real challenge lies in supporting diverse model architectures (e.g., BERT, Qwen, Colbert) with different attention mechanisms and positional embeddings. The engine re-implements forward passes with variable-length FlashAttention and handles model swapping via a least recently used eviction policy. Makraduli also explains that context management for agents requires small models to pre-process data, referencing Andrej Karpathy’s graph-based knowledge bases and Chroma’s own model. The talk details the infrastructure layer including routing, auto-scaling with Prometheus, and GPU provisioning using spot instances, all open-sourced as SAI (Superlinked Inference Engine) with Helm charts and Docker images.

Continuous Profiling for GPUs — Matthias Loibl, Polar Signals
Jul 22, 2025 · 11:31
Matthias Loibl of Polar Signals explains how continuous profiling for GPUs maximizes GPU efficiency using low-overhead, always-on sampling via eBPF. He contrasts tracing (high cost) with sampled profiling (e.g., 100 Hz, <1% overhead) and details GPU metrics collected from NVIDIA NVMe, including utilization, memory, clock speed, power, temperature, and PCIe throughput. The platform correlates these with CPU stack traces to identify bottlenecks, such as Python and CUDA functions underutilizing the GPU. A new GPU time profiling feature records the duration of CUDA kernel executions, showing actual time spent by functions on the GPU. Deployment runs on Linux with a binary, Docker, or Kubernetes DaemonSet; early adopters like TurboPuffer use it to optimize their vector engine.
Powered by PodHood