Web scale0:00
So, hi everyone. Thank you so much for taking the time to join this session. I hope I'll—or at least I can guarantee I'll do whatever it takes to make it worth your time. My name is Omer, I lead the product marketing team over at Bright Data.
Just by maybe a quick show of hands, who here is familiar with Bright Data? Okay, we can do better. I'll pass it on to our brand team. Bright Data is a web data company. Basically, we help more than 20,000 teams around the world, including more than 70% of the world's biggest AI labs, to extract data from the web.
Just to put this in perspective of what scale we're talking about, we're talking well over 50 billion pages of HTMLs every day, more than 20 petabytes of video, audio, and other media data. So that's just the perspective of the type of work that we do at Bright Data.
But enough about us. Personally, I joined Bright Data about 3 years ago, which essentially gave me front-row seats at everything around AI and the web and how they started to actually connect. It sounds very old, but if you think about it, only maybe less than 2 years ago,right, we were able to start using Access the Web and Search the Web through Claude or through ChatGPT.
That option didn't even exist in the earlier versions,right? So that connection, that way in which both AI and the web are starting to converge, is something that is still evolving and evolving rapidly. And that's part of what I want to try and shed some light on today, and talk about a new emerging breed of companies that's coming out of this connection.
I think we can basically agree that the web is by far the world's greatest source of data. At least historically, when it comes to Bright Data, that's all we cared about,right? It's helping our customers extract data from the web.
But with the emergence of AI, and the emergence more recently of AI agents that need to do knowledge work,right, the web is no longer just a source of data. We can actually start looking at it as a source of context.
Context in the sense that if I do knowledge work and I have knowledge agents that support my work, I want to go out to the web, find the information that I need, use it as context, but keep on working.
So the data itself is only a step in the process for something bigger, for the actions I need to take, for the conclusions I need to draw, and for every downstream application that follows. The first one's to figure this one out.
Oh, sorry, even before that. But one thing that I want all of us to bear in mind, because this is going to follow us through the rest of this conversation: the web is messy, it's unstructured, and most importantly, it changes all the time.
Data decay2:43
This is a chart that shows data decay,right? It's an analysis done by our team that shows data decay. Basically, how long after a new page, a new piece of content goes live, it is no longer relevant,right? So social media, it's easy for us to understand, it's far less than a day.
But also news, finance, retail. Thirty days later, data that was collected is mostly no longer relevant. And when we acknowledge that, this simple notion, we understand that extracting context from the web or relying on the web is not a snapshot.
It's not a one-time effort. It's not even a monthly effort. It's something that we need to keep on doing. It's something that we need to look at as an ongoing process, and something that we need to be mindful of.
The first ones to figure it out were, of course, search companies,right? Only, what, 3 years ago we were on the far left,right? This is in our lifetime,right? Three years ago we were on the far left. Everything was Google.
Search fragments3:37
There was complete and total dominance up until 3 years ago. For the past 20 or so years,right, that's what we're talking about. Purely for humans: search something, go, collect the information you need, and carry on. Then fast forward maybe one and a half years ago, 2 years ago, search began to appear within the LLMs, within the chatbots,right?
Which already started blurring the line between humans and agents, because now the same bots also have the same web search available. So the same LLMs have the same access through API for the bots. So we started seeing that convergence happening,right?
And so for the first time, we're seeing more and more traffic flowing down, search traffic, search intent flowing down these channels, not only to Google. And last but not least,right, we now see a whole breed of companies, the AI search companies.
If you're familiar with them, I caught a talk yesterday by Will, the CEO of Xa, and Parallel, and you.com, and Tavili, and a bunch of others. They are purely built and indexing the web, especially for agents. They're not even looking at the humans involved anymore.
So Google's dominance when it comes to search, if Google was synonymous with web search, that is very much shaking. And when there's blood in the water, the sharks come. Just last week, Amazon announced I don't know how many of you saw it that they developed their own index and started allowing the ability to retrieve data from the web, to retrieve context for agents on AgentCore.
Amazon developed their own search engine.
Two weeks before that, it was Microsoft. Microsoft always had skin in the game, man. One or two percent of the world search traffic went to Microsoft. But they have now repackaged it and launched it again as part of Web IQ, as part of their suite for agentic development and orchestration.
So we're seeing more and more this space becoming crowded. But when we're talking about context, and we're looking at this through the lens of search, I believe it only tells us part of the story,right? I can search for, I don't know, what's the cost of a certain pair of sneakers this morning,right, on a certain website.
Search limits5:42
I cannot really search for how has that price changed over the last six months, what discounts it had,right? I can search for what open job positions we have at Bright Data. We do. I urge you to go have a look.
But I can't see how that was a chart and how that changed over time and how the headcount of the company changed over time. And all of that information existed on the web simply back then. So when we actually start to think about it, we understand that there's much more context in the web than what web search allows us to extract.
And this is what we started seeing in recent years, a whole new breed of companies rising. We like to call them internally CaaS, Context as a Service, because that's what they do. They allow agents to tap into them, MCP, CLI, just pure good old API, and actually start extracting data to retrieve data so they can reason over for whatever knowledge work they are responsible for.
CaaS6:32
We see this happening in e-commerce, we see this happening in travel, we see this happening in finance, in market research, in HR, in real estate, in a bunch of other domains. I'll show a few examples in a second,right?
What all of these have in common is that they don't only just discover the web, you know, in terms of think crawling, think searching, think all of that, accessing, extracting the data, and indexing it. They take it a step further.
They actually develop knowledge graphs to start structuring all of the entities and to dedupe them. And they start enriching them with a lot of different sources. So they actually start merging all of that data. If you think about it, they kind of behave like vertical search engines,right?
They are very, very, very good search engines for something very specific. And it's already in full motion,right? So as I said, we see this in finance and in market research and retail and e-commerce and GTM and sales intelligence.
What all of these companies, by the way, have in common, they're all part of Bright Data's startup program. If you are a builder, and this is a hot space to go in because I think we're only tapping the surface, I invite you to scan this and apply.
Up to $20,000 in credits and all sorts of co-marketing. But that's enough self-promotion.
So CaaS, as an industry, is already in full bloom. And as always with these situations,right, also the traditional players aren't left too much behind. They're at the bottom. You see good old data as a service. You see ZoomInfo,right?
By researching for this presentation today, I also saw that they launched that thing at the top. It's called GTM.ai. You can only imagine how much they paid for that domain. But they launched it as a secondary brand for ZoomInfo that is catering specifically for the need of agents.
Look at the wording,right? They talk about GTM work,right? That knowledge work, that research that you do when you need to prospect, when you need to do headhunting, when whatever it is you need to do that involves people mostly, straight from Claude Code, straight from Codex, or any other agent.
They understand the gap,right? So yeah, I say it's funny to think of CaaS as an evolution of that. And it is, in a way, but it's catering for a very specific need. As much as the AI search engines are different than Google, when agents need them, it's different than people.
When we let this one sink, that at the very least we have two different types of paths to complete knowledge work as an agent, we can start thinking about this in terms of web context engineering. We can start thinking about this in terms of how do I optimize for the specific task.
More importantly, when things come as it is, how do I optimize this for breeds of tasks? How do I do this for various parts of the organization that I'm building for? If I'm an AI engineer, I need to serve different teams.
They may have different needs. It's very tempting to throw AI search at all of them, but maybe that's not optimal. Maybe I need a combination of both. Maybe I can start seeing all sorts of cost efficiencies emerge from that.
So for the second half of this presentation, we actually went ahead and created a test. This is not a benchmark. You won't see any something concrete that I can say with great confidence other than the actual research that we did, because we wanted to start unraveling the different considerations and how do these two stack up against each other.
The test10:01
So we designed a test. We went for something basic. We said, OK, let's take a company, an entity, and try and enrich it across 25 different fields. Some of them are very easy, you know, the company domain, the name, the headquarters.
But some are more challenging,right? Things about hiring and people and something. And we built a simple agent, a loop in a loop that knows that uses Opus 4.0 as the harness. And it starts to go over field by field, go out, search for it, or retrieve it from the CaaS, do it again and again and again until it completes and brings back, sets some guardrails, you know, like budget and stuff, just to keep it fair.
And I'm going to share with you the results. So the first thing that we would care about,right, being knowledge work, would be sorry, we ran it 100 times on all of the sponsors of today's event. So the first thing that we saw in terms of coverage is that there's pretty good convergence.
Coverage11:04
They all did fairly well,right? I'll get to the two at the bottom in a second. So search were consistent performance. One of the major CaaS providers were also very well. The third one, by the way, you can see on Locker and SERP.
SERP is good old data, good old Google. We basically did the same thing, just with Google, and it performed pretty well in extracting that information. Native is Claude's own search. And you see that they converge really well. I was originally surprised about the two CaaS solutions at the bottom.
It was counterintuitive. I expected CaaS to dominate this thing, because that's you had one job,right, to map out these companies. But after diving into it a bit more, you understand that, well, they are limited in the sense that they know what they have about an entity.
If I asked it a question that is beyond that, they will never have that data,right? Unlike a search, it can go out and continue searching and exploring it. If they didn't collect data about the recent job hiring, it will never be there,right?
So it makes sense that they are a bit behind, but I'm sure at the same time that they have a lot of other advantages that we simply didn't ask for, a lot of other fields that they didn't have that aren't represented.
So again, it creates some complexities when how do we measure coverage when it relates to the specific job that we need to do rather than in general. The second thing we looked at was cost, of course. Here we started seeing it spread out a bit.
So you can see that massive bulk in the center. Most of the search and the CaaS and even using Google,right, converged to pretty much the same cost, only different,right? The CaaS was just about the service itself, what you pay the vendor,right?
Cost12:32
All of the other search solutions, you also needed a lot of token burn to actually structure that data so you can actually act on it and use it as something retrievable,right? So it's the same output. Native, obscenely expensive.
And the CaaS on theright, I'm sure you're all familiar with, they're by far the most expensive in the industry. I will not name and shame them. Interesting, you see that small CaaS there at the left, that CaaS number two, that were very cheap.
They're also the ones that are here at the bottom, which is funny because what I believe is happening there is that we're seeing even within this industry, niche players that have lower quality data but much cheaper, they're already carving that niche of the long tail,right, of small shops or small usage that don't want to pay as much and don't need as much data.
And we're all seeing them branch out there.
Most of you here, I presume, are engineers. So there's a very evident question that we did not ask here, which is, what is the one thing that an engineer would care about?
Thank you. Let's talk about scale.
This is the cost, not for the whole hundred. This is the cost per one, for one record.
What happens if we need a million? Now, yes, a million records will not fit in a context window, obviously. We're not talking about a single run that needs a million. You can think about a million in terms of the frequency,right?
If I am a market researcher, I do due diligence for private equity. I revisit these companies all the time. I ask more questions about them as the time goes by. Was there any new news about them? Was there anything that changed?
Did somebody join? Did somebody leave? Do they have new hires? I keep on asking the same thing. So when I'm talking about this, the multiply by a million, it's not just about the number of companies. It's the frequency in which I'm asking it.
Frequency is the cost killer when we talk about these. And we need to acknowledge that,right? We're thinking about this in terms of web context engineering. We're starting to look at it differently. Every repeated query costs the same as the first, even if it brought back the exact same answers.
Nothing changed, pay up,right? Noise false positives for sure go in. Token costs,right? We saw that the model, the very high token, we know that doesn't shrink well over time. There's always some volume element in terms of the cost, but it's not the same as flatlining,right?
And if we bring this back to knowledge work, this is where we see teams that are starting to cut corners. So I won't research this company every day. I'll look at it once a week or once a month.
Cutting corners15:22
I won't ask that question now. I don't want all the results. I'll only take 10 results, 20 results, something. So we already have the setup. We have what we need to do the knowledge work. But at the same time, we're not extracting all of the value because we're starting to be conscious about cost,right?
Basically, we're renting context. We're not owning the context that we use. That is a very important distinction. Again, if we're good engineers and we ask ourselves, what about scale? The second and most obvious thing that will come to mind now, so how about we build it?
What if we take all of that web data ourselves and stick it in some vector database and try and see what comes out of it? So I asked my AI engineer to do exactly that. Again, this is a test.
Build it yourself15:59
This is not a benchmark or full-blown operation. This is a day's work at best just to illustrate the concept and to show something about the cost efficiencies that you can generate potentially by doing it yourself, potentially, in specific scenarios.
The test, simple. Take the company name, nothing but run it through Google, find the relevant entries, the relevant URLs of that company in various websites that have all of that information,right? You use search when you don't know the source.
But when we're talking about company enrichment, we all know these sources. We all know where that data comes from. The CaaS also bring it from them. ZoomInfo bring it from them. It's the same thing over and over again.
Why not just go straight to the source? Why are we doing that middleman thing? LinkedIn companies, LinkedIn jobs, Crunchbase. Right there, we have scrapers for those. You just tap in and you start paying as it pay as you go.
We built two dedicated scrapers. We have a new AI tool called Scraper Studio. It basically lets you build a scraper for any website in less than five minutes, all powered by AI. And then it also has a self-healing function,right?
So if the website changes, it fixes itself and keeps on going. Merge it all into one entity, basic heuristics. If there is conflict, choose that over that. And eventually, we have a data set of these 100 companies, zero AI cost involved.
There's no tokens. Coverage, fairly well. Not amazing, not the best that we saw here, but stacking up pretty well. And again, this is just a day's experiment. Probably not even as much, OK? So again and again, very specific tasks, very limited context, very limited situation.
Tread lightly and proceed with caution when it comes to conclusions. The real story is not this. The real story is this.
That's what it cost, to just go and fetch that data that is out there. We think about knowledge graphs. We think about entities. But if you think about, for example, LinkedIn, the data is already structured in form of entities.
There's an entity for a company. There's an entity for a person. There's an entity for a job, and they're connected between them. Sometimes the ontology is already there. Again, this is not the most complicated of scenarios, but this is pretty damn good.
Now,
yes, it took time to set up. So it's not really fair to compare apples to apples when it comes to the cost, because these are out of the box. You can just tap into the API. That one that I just showed you required some setup.
Tipping point18:34
Let's say it's a week. Let's price it at $5,000 just to give us some perspective. We can actually start thinking about this in terms of a tipping point. We can actually start thinking about what is that tipping point in which it makes more sense for me to build it myself,right, than keep on renting it.
Now again, everything to the left of that dot, in this case, it was just over 15,000 entities or queries,right, when we think about it. So it made sense to do it at this point. Maybe it's not 15. Maybe it's 30.
Maybe it's 100,000. Maybe it's 10,000. It really depends on the use case. But there is a tipping point in which it actually makes sense to do it yourself, which leads us to the fact that both AI search and CaaS and all of these solutions, they're very good in the sense that you can just plug and play.
But if your knowledge work needs,right, are persistent and consistent and to a certain degree may even continue escalating and growing, then this is perhaps a direction to start considering. Maybe I can just go ahead and build my own, because the nice thing about it is that all of the things that we see on the left, up until the third part, is upfront investment.
And the most important thing, that whatever retrieval happens later on from the agents,right, is free. Not really free, but you get what I mean,right? There's no added cost. I can just ask that question over and over again. I did not like the first answer.
I'll ask it again. I'll ask it 100 times until I get what I need. I have no more fear, no more cutting corners, which is maybe the most important thing. And I'm leaving aside the fact that this is also custom business logic.
I can connect it with my own data. There's all sorts of other advantages of owning it. You know, we'll keep it to the imagination.
Remember, we asked about a million, not about 15,000. This compounds. This compounds greatly,right? We need that horizon. Remember, the web keeps changing. We saw the staleness of the data and how the data decays. So we need to be thinking about this in the long run and how this will evolve when we keep on asking the questions about the entities that we care about.
Just to wrap it up, so AI search, CaaS, they can get you very far when what you need is ad hoc and what you need is always changing. When sometimes you look at different things, even the mix and match of them for certain tasks, use this.
Wrap-up20:51
For certain tasks, use that. I'm sure, again, I just showed a test. There's a lot of ways to optimize it, just like any other context engineering, and use lighter models and use other stuff. There's a lot of great stuff to be done.
But eventually, the frequency will come and bite you in the ass when it comes to cost. And that's something to be mindful of. And there's a fair chance that that tipping point is much lower than you think. And that's something that as we design these systems, when we think about web context engineering, we need to be mindful of that.
And last but not least, the last slide we showed, owned context compounds while rented decays,right? It's not a one-time task. Again, if it's a one-time question, use AI search. It will be amazing. When you need to do it over and over again, there's a fair chance that it will not that you're missing out on potential compounding effect and you are losing out.
Thank you very much.





