# Agentic Sites: Building Hyper Personalized Websites — Carlos Sanchez, Adobe

AI Engineer · 2026-08-29

<https://aiengineer.podhood.com/20276bb2-6930-4e00-995b-2d90bbcc3beb>

Carlos Sanchez, principal scientist at Adobe, argues that agentic sites can deliver hyper-personalized web pages in real time by treating the whole site as a corpus and generating only specific blocks. He demonstrates a coffee-machine site that builds a personalized camping page in 1.64 seconds using Google's Gemma 4 on Cerebras at 2,300 tokens per second. Model choice must be evaluated per site for accuracy and speed, he says, showing 1.1 seconds against 4.6 for the runner-up, and this doesn't need a frontier model because work is choosing and arranging blocks. Marketers define personas in natural language, browsing signals feed an 'audience of one' loop for pre-generated recommendations, and a tool builds an agentic site for any URL in less than an hour.

## Questions this episode answers

### What is an agentic site, according to Carlos Sanchez?

Carlos Sanchez defines agentic sites as sites that look at the intent of the browsing user—what the user is doing and trying to achieve—and personalize pages in real time for that current user. The end goal is to drive higher engagement or conversions, depending on what the marketing team wants to achieve.

[1:02](https://aiengineer.podhood.com/20276bb2-6930-4e00-995b-2d90bbcc3beb?t=62000)

### How does Adobe's agentic site approach avoid hallucinations on branded pages?

Carlos Sanchez says marketers have very strict brand guidelines, so the whole site is not generated. Only different blocks are customized by persona, and the entire site is used as a corpus through retrieval, grounding whatever is generated in the existing site content and preventing hallucinated copy.

[2:34](https://aiengineer.podhood.com/20276bb2-6930-4e00-995b-2d90bbcc3beb?t=154000)

### Why does Carlos Sanchez use Cerebras for generating agentic sites?

Carlos Sanchez evaluates models per site with Promfu, looking at both accuracy and speed because a page that takes more than one or two seconds loses users. On an example with 15 prompts, Cerebras running Gemma 4 achieved an average latency of 1.1 seconds for a generated page, compared to 4.6 seconds for the next option, so it was chosen.

[6:43](https://aiengineer.podhood.com/20276bb2-6930-4e00-995b-2d90bbcc3beb?t=403000)

### What does Carlos Sanchez mean by 'audience of one'?

Carlos Sanchez calls the project 'audience of one' because in marketing people have always dreamed of personalizing things for each individual. His demo shows a fully generated coffee-machinery site where browsing signals are bucketed and the site generates a custom page, including a For You page, for that specific visitor.

[12:34](https://aiengineer.podhood.com/20276bb2-6930-4e00-995b-2d90bbcc3beb?t=754000)

## Key moments

- **[0:00] Agentic sites**
  - [0:14] Agentic sites build hyper-personalized websites in real time, and Carlos Sanchez promises a live demo of what Adobe is building.
  - [1:02] Agentic sites infer what a visitor is trying to achieve and personalize the page in real time to drive engagement and conversions, says Carlos Sanchez.
- **[2:20] Block personalization**
  - [2:34] Agentic sites don't generate whole pages: they personalize blocks and use the entire site as a RAG corpus to stay on-brand, says Carlos Sanchez.
  - [4:30] Model choice for agentic sites is site-specific, so Adobe continuously evaluates LLM providers on accuracy and speed, says Carlos Sanchez.
- **[4:37] Architecture**
- **[5:46] Model evaluation**
  - [5:58] Agentic page generation must finish in 1–2 seconds because faster sites convert better, Carlos Sanchez tells the AI Engineer audience.
  - [6:43] Cerebras with Gemma 4 generated agentic pages in 1.1 seconds average vs 4.6 seconds for the runner-up, Carlos Sanchez reports.
- **[6:58] Speed results**
  - [7:54] Agentic sites don't need a frontier LLM — the work is generating text and arranging blocks, says Carlos Sanchez.
- **[8:07] Frontier models**
  - [9:04] A 'For You' recommendation page can be pre-generated as users browse so product picks are ready, says Carlos Sanchez.
- **[9:16] Pre-generation**
- **[10:23] Personas**
  - [10:32] Agentic sites can reorder blocks and generate images on the fly, says Carlos Sanchez, pointing to the new Nano Banana Lite model.
  - [12:34] Adobe calls hyper-personalized sites 'audience of one' — marketers' dream of tailoring pages to each individual, says Carlos Sanchez.
- **[12:44] Audience of one**
  - [13:09] Live demo: browsing signals bucket a visitor as 'exploring' and trigger a generated For You page on Adobe's coffee site.
- **[13:53] Live demo**
  - [14:22] Live demo: asking an agentic site for a coffee machine to prepare coffee while camping instantly generates a personalized page with recommendations.
  - [15:30] A live agentic query generated a full page in 1.64 seconds, including a round trip at 2,300 tokens/sec on Cerebras Gemma 4.
- **[16:20] Tooling**
  - [17:02] Adobe's tool builds an agentic site for any URL in under an hour — Carlos Sanchez generated one for the AI Engineer site with side-by-side conference comparisons.
- **[18:18] Future vision**
  - [18:18] Agentic sites can answer voice queries on Google TV and generate a personalized buying page, Carlos Sanchez envisions.
  - [19:30] Agentic personalization is possible now and will only get better, cheaper, and faster, predicts Carlos Sanchez.

## Speakers

- **Carlos Sanchez** (guest)

## Topics

Context Engineering, Content Engineering

## Mentioned

Adobe (company), Cerebras (company), Cloudflare (company), Google (company), Adobe Experience Manager (product), Bedrock (product), Gemma (product), Nano Banana Lite (product), Promfu (product)

## Transcript

### Agentic sites

**Carlos Sanchez** [0:14]
Hello. Thank you for coming. Um, I'm going to talk to you about agentic sites, how we call it, is bu- building hyper-personalized websites. I'm not going to just talk about it, I'm going to show you what we're building.

Um, I've been working on, on this project for, for a bit now, and we'll try to show you what is possible today with, with AI. Uh, I work at Adobe at, uh, I'm a principal science at a product that not many people know, Adobe Experience Manager, content management.

We run a lot of, uh, websites, properties for big brands, and my background is in, in open source, uh, contributing to, to a lot of foundations and projects.

What are agentic sites, and how are we building this thing? So we are looking for sites that are, uh, looking at the what intent the user browsing, uh, has, what is the user doing, what is the user trying to achieve.

And the end goal is to personalize these pages for the, for the current user browsing so that eventually this, uh, drives, uh, higher engagement or, uh, conversions, whatever the marketing teams want to, want to achieve. And these pages are personalized in real time based on the, on the user that is, uh, accessing the site and what is the, what is the user doing.

The stack we're using is AMH Delivery, so this is the part of the product we, we have, uh, where all the content is, uh, on the edge. And then we have, uh, backend service that powers this experience with, uh, uh, different LLM providers, uh, LLM services.

Uh, we use Cerebras for fast inference, or we can use also we tried Bedrock and a, and a bunch of others. I'll be showing Cerebras today, um, and you will see the reason why. The, the engine that is personalizing this, uh, this it's, it's, um, it's using the rich con-content and blocks.

### Block personalization

**Carlos Sanchez** [2:34]
So different blocks on the site are customized depending on, on what the user persona is. We don't want the, the whole site to be generated. I mean, if you talk to marketing people, they, they have a very strict brand guidelines.

You don't want to just come up with or have some hallucinations there. So the what is personalized is different sections of the site, and we use the whole site as a, as corpus. We build a rack from the whole site, so what is generated is, uh, grounded on, on the existing site.

We try to solve the problem where, yeah, one size fits all. We want hyper-personalized experiences. Also, we want to help, uh, our customers to do more automatic, um, authoring, so not having to create thousands of different variations of the site, but, uh, use AI for this, and then do these multiple layers of, of personalization.

Uh, some examples of, of what we are doing or are showing the demo is, uh, per instant persona ad-adaptation, query generation when the user searches for something on the site, the page with the results is customized for them, and also, uh, something like recommendations where after you browse the site for a period of time, we recom we can create a page that recommends something based on, on, on what you are what we think you are looking for.

For marketers, uh, they can define this strategy on natural language, and they can use analytics to, to drive the loop of personalization and what is the end goal and how this goes back again to change to adapt the personalization to improve that, uh, whole cycle.

Everybody's talking about loops in this conference, so that's, that's one of the loops there.

How the architecture look like. So it's a dynamic front end with some blocks, what I mentioned before, and with, uh, edge delivery services is basically you compose these blocks, and, uh, they are updated on, on real time through with the AI.

### Architecture

**Carlos Sanchez** [4:48]
The backend, uh, we, we do the, um, evaluation of the models and the providers. And one thing we realize is, is that this is very dependent on the site. So we have a bunch of prompts, and we look, uh, we run it across a huge variety of, uh, models and providers, and then we look at the accuracy, we look at the speed, but this is going to depend highly on what type of site, like how big is the site, how, I don't know, what different, um, what different, um, area is the site, uh, targeting, what, what type of commerce it is, and so on.

So we, we run this, this, um, evaluation continuously. We use, uh, Promfu. Uh, anybody heard about Promfu? Okay. Some people. So Promfu allows you to evaluate models, um, prompts against, uh, multiple models, providers, and, uh, you can do local models and any of the a bunch of, uh, OpenAI compatible, uh, providers and, and, um, a lot of them, basically.

### Model evaluation

**Carlos Sanchez** [5:58]
We look for two things. Why? Accuracy. That's, that's typically what people look for. But also, we want the speed because we don't want the site generation to take more than one or two seconds,right? Because people, um, this is already, uh, proven that people want the, the faster the site, the more conversions it, it generates, or the, the better the experience it is for the user.

Um, yeah, what I mentioned is different sites may have different requirements, uh, so you may have to run this, uh, evaluation of models depending on the site. This is, uh, an ex uh, we, we executed this, some of these queries.

So we have, uh, 15 prompts for this example site, and we have, uh, at the top, you can see with Cerebras on the Gemma 4 model that was announced last, last week, we can get, uh, an average latency of 1.1 seconds generating a page.

### Speed results

**Carlos Sanchez** [7:01]
You, you can compare that to the second one, which is 4.6 seconds,right? So the difference is huge. And that's why, uh, we use Cerebras for, for this use case. And, uh, you can see that different providers, different models have different, um, different speeds.

And here is, uh, let me I can show you the whole thing here. Not this one. This one. Right. So at the, at the bottom, we have other, other that sometimes, uh, maybe some of them may be good.

They don't need to be perfect, but they're good enough if they're fast enough. So that's going to be the, the kind of decisions that you need to make on whether the model is good enough for your, your, your use case or not.

Yeah. We're looking yeah. Average 1.1 seconds, and then the, the next ones are going from 4 seconds, uh, higher. And you don't need a huge LLM to do this sort of work because you are generating text, uh, you are deciding where to put blocks and how to organize the website.

### Frontier models

**Carlos Sanchez** [8:12]
You don't need lots of information for that.

So this browsing and the queries, uh, is are being recorded. So these are the metrics or the, um, the, the data we gather from the user, and this is fed into the LLM to personalize the site. And then in this example, we personalize the hero card, the products, the block feeds, and, and the navigation based, based on the persona.

Also, what are, uh, some of the buttons like our call to action navigation, you can also we can also personalize those. We, uh, we create and I'll show you the, uh, For You page, which is a recommendation. And this is an interesting one because this you could, uh, pre-generate,right?

As the user browses your site, you gather these signals, and you could keep generating this. So in this case, you wouldn't need so such a big speed. But, um, but that's interesting because it, it would be if a user wanted to buy something, you could just say, "Okay.

### Pre-generation

**Carlos Sanchez** [9:23]
For you, I will recommend these three products," or, or something like that. Um, yeah. And then they can see this recommendation, and if they go there, that, that could be pre-fetched for them. And obviously, you have to keep updating it as the user navigates around the site and, and so on.

So that, that's also something to consider on the cost, cost, uh, implications of doing multiple generations, multiple LLM calls.

When, uh, when the user runs a query, a dynamic personalized page is shown to them. Uh, when the these queries are also grouped into personas or intent types. So what is this guy what is this guy trying to do in the site?

He's trying to buy something. He's trying to just get information. So you can get marketers to decide what type of groups, how many groups you want to have, how you want to deal with, with customers. And the AI will choose the, the blocks and the suggestions for, for those groups of people.

### Personas

**Carlos Sanchez** [10:32]
Um, and we can adapt, yes, the, the different blocks, the, the, the sequence of the blocks, and, uh, media. You could also do media. One of the things we consider is, uh, there was some a model announced, uh, today or yesterday, the, the Nano Banana Lite.

So you could even generate images very fast on the fly. Obviously, not as fast as TESTS, but, uh, that's also, uh, something that would be I don't I don't know if is that something like marketing people would want to have generated images.

That depends on, on the quality a lot if it's on brand. And the site, uh, in this example, we have a, a product site, and then we have guides, experiences, blocks, and the whole response of the LLM is grounded there.

And, uh, there's comparisons. We can do comparisons between products that are tailored, and the product pages can be tailored for the, for the user.

Okay. This, this is a bit of the stack. Um, not going to spend too much time here, but the browser, you have some layers. You have the browser where the signals get, um, get, uh, uh, from are g are got, uh, are got from the, from the user.

And then we have the backend. Uh, we can have the backend. We run this, some of this in, in Google. Some of these are in, are Cloudflare. So the backend is basically just calling the LLM and doing some reasoning using the rack that is built on, on the site to do the generation.

And you have obviously, you have to have the vector database, the inference, uh, machinery, and, uh, the Adobe Experience Manager is doing the serving the, the at, at the edge is serving the, the pages and the static content.

So let me show you because I think this is, uh so we call this, uh, audience of one because the idea of, um, in marketing, uh, they, they always dream on being able to personalize things for each individual.

### Audience of one

**Carlos Sanchez** [12:49]
So we call it, yeah, a, audience of one. So I have this, this site. Uh, this is a site that is absolutely generated. Uh, example site is a coffee, uh, machinery. So I can go and, and read some stories, and I can go and look at some products.

Let's go and look at this product. I can spend some time here.

Uh, let's go and click here. Okay. So I'm loo I'm browsing around the site, and I have this debugging tool thing, uh, which, uh, let me go here, I think.

Let's see. So down there is the signals that the, that the browsing, uh, is giving us. So I don't know if you can see it much because I cannot see it much. Then so the user is, uh, bucketed into the exploring category.

### Live demo

**Carlos Sanchez** [13:57]
We have the pages that, uh, have vis uh, have visited, and then we have, uh, how much time is spending on each page. All of this data is now available for the LLM. So if I go here, I already have a For You page that was generated for me and, uh, based on my browsing.

And you will not notice that it's s-slightly different than everything else. But if I go here and I run a query, uh, like I want I am looking for a coffee machine to, uh, prepare coffee while camping.

The site is this was just generated for me. And then you're going to see some things like the text is customized. Camping shouldn't mean compromising on your, uh, whatever routine. Uh, the coffee tips for camping, um, machine, uh, that are being recommended, the Arco Viaggio, and, um, or the Nano, which are, um, good for, um, for the, um, for a camping trip,right?

So you saw how fast this was. I'm going to run it here, uh, something similar that I have here, and I can run it on the debug mode here, and you will see let's make this bigger.

Total time 1:64 seconds to generate the page. So this includes a round trip to the LLM. This is using Cerebras Gemma 4. So the, the Gemma model from Google running on Cerebras on their, uh, very fast chips, uh, we get 2,300 tokens per second, which is not bad, I would say.

And if I run it again, uh, probably something like that, uh, the LLM time is 1 second, and again, 2,200 tokens per second. This is something that we only dreamed about before.

On the on this site, example site, we have some other options. Uh, so because we've, we've been showing this to customers, so we have the, the ability to change, uh, the different models, tempera temperature tokens, and so on.

### Tooling

**Carlos Sanchez** [16:26]
And we can, uh, we can show, uh, and try the different models and see how they behave. Besides automatic test with Promfu, then we can, uh, manually come and, and tweak things and see and see how that how that works.

And, uh, we also have, uh, Of1 Labs. So we have we built this tool that generates an agentic site for any site we want. So if somebody wants to have a, a demo for a customer, come here and enter the URL.

In less than an hour, you have an agentic site. I did this last week with the AI Engineering site, and I got this site that is just a search box and a few things. And, uh, let me open it here.

The full page. Not this one. Yeah. Okay. So I could say, uh, Europe AI conferences. So these suggestions are also AI generated, and I get a page that is, uh, more focused on it should be more focused on, on the on this European conferences.

If I go back did it go? I can search for anything the same way I did with, with the Arco. So I as a specific there was someone that was generating a good comparison side to side. Let me see if this one.

Okay. Here. This one. I went and this generated a page with, uh, very good comparison. If I'm looking at two conferences and I need to decide, if I figure out that the user wants to do that, this is great because that gives them a side by side comparison on the fly.

Now,

this, this is I think this is cool already, but then we have, uh I have this idea that probably the, um, a bunch of people are we are talking about is the web dead? Is, is the web the future still?

### Future vision

**Carlos Sanchez** [18:32]
And so on. Nobody knows. But we can also do something with this, uh, with this audience of one, this, uh, generative sites. So imagine you have, uh, you have your personal assistant, and you ask a query through in this case, through Google, and you say, "I want to buy " I don't remember what the query said.

It was something like, "I want to buy, uh, a machine," and I get this on my Google TV,right? So this is absolutely personalized to my query. Okay. Nope. Go back.

This is absolutely personalized to my query. So I'm there in my living room. I don't need a phone. I don't need a computer. I don't need anything. Just my voice and something that will, uh, kind of show me something that is absolutely personalized to, to me.

Okay. So that one. So what I was trying to show, and hopefully, you remember from this session, is that this is now possible. It's only going to get better from here on. It's only going to get cheaper. It's only going to get faster.

And you will be able to have, uh, huge personalization options for sites and for other things. And you can do this, uh, with intent-driven. So what is the what is my user trying to do? What does my user want to buy?

These sort of questions. And you can, uh, assemble a page just for them. And you can also do this with, uh, multiple models. And, and eventually, it's just going to be faster and faster,right? So that's it. And thank you for coming, and I hope you, you got the idea.

Thanks.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
