Episodes from AI Engineer about Image Generation.

Training Krea 2: What matters in generative model training — Sangwu Lee, Krea.ai
Aug 18, 2026 · 21:46
Sangwu Lee of Krea.ai explains what went into training Krea 2, its open-sourced image foundation model, arguing that once architecture is locked, data is everything. To keep stylistic diversity over the 'most boring average person' consistency of production models, Krea filters billions of images: no AI images, OCR and vision-language captions, hash and embedding dedup, distilled VLM classifiers, sparse autoencoders as unsupervised taggers for watermarks and borders. World knowledge is checked against Wikipedia concepts by PageRank. Training goes from 256 to 1K resolution through pre-training, mid-training, SFT, preference optimization, and GRPO-style RL, plus a prompt expander. Lee wants a single clean transformer without VAEs/text encoders and VLM-generated bounding boxes or scene graphs.

"The biggest challenge in your stack? Evals, Evals, Evals" - 2026 State of AI Engineering results
Jul 21, 2026 · 19:47
In the 2026 State of AI Engineering survey presented by Amplify Partners' Barr Yaron, 1,048 AI engineers reveal that cost is now a first-class engineering constraint—40% say it regularly shapes how ambitiously they use AI. Agents have exploded: 95% of teams now use agents, and 89% of those agents have write access, tripling from last year. Image generation adoption doubled to 36%, while audio shows the strongest intent-to-adopt at 56%. Open-weight models augment rather than replace closed models—45% use open-weight, but over 90% of them also use closed models. Evals remain the top infrastructure challenge, and inference is the most bought layer, while prompt management (61% built in-house) stays close to product logic. Teams report 97% net positive impact, but 59% fear long-term liabilities from AI code, and over a third say non-developers now ship features.

Any-to-Any: Building Native Multimodal Agents - Patrick Löber, Google DeepMind
May 20, 2026 · 16:21
Patrick Löber, a member of the technical staff at Google DeepMind, explains how to build native multimodal agents using the Gemini API ecosystem, covering multimodal understanding, native image and speech generation, and real-time interaction via the Live API. He demonstrates constructing a NotebookLM clone as an agentic system where a reasoning Gemini model decides whether to generate an infographic or podcast-style audio using function calls to specialized models like Nano Banana for images and a text-to-speech model for speech. The episode details practical implementation: uploading PDFs, video, and audio files; using context caching to reduce costs by 90%; and generating infographics or multi-speaker audio directly from prompts. Löber highlights that native generation models understand world context—like drawing arrows on a map to produce the Golden Gate Bridge—and that the Live API enables audio-to-audio interactions with a single architecture, supporting multiple languages and accents.

FLUX, Open Research, and the Future of Visual AI — Stephen Batifol, Black Forest Labs
May 8, 2026 · 22:32
Black Forest Labs (BFL), the team behind Stable Diffusion and Latent Diffusion, has released a series of open FLUX image models pushing toward visual intelligence. After FLUX.1 and the first open-source editing model FLUX Kontext (7–8 second edits), FLUX.2 achieved state-of-the-art text-to-image and multi-reference editing, and FLUX.2 Klein generates and edits in 300–500 milliseconds for near real-time use. BFL also published Self-Flow, a scalable self-supervised approach that trains multimodal models across images, video, audio, and actions without external encoders, outperforming baselines in all modalities. The episode explains how Self-Flow reduces artifacts (e.g., corrects text rendering and anatomy) and converges 70× faster with representation alignment. BFL’s roadmap includes world models that simulate geometry and interaction, aiming to train agents for robotics and automation.
Powered by PodHood