AIAI EngineerAug 26, 2026· 15:04

The Missing Layer in Agentic AI — Giedrius Šteimantas, Oxylabs

Giedrius Šteimantas of Oxylabs argues the missing layer in agentic AI is web scraping infrastructure, applying ten years of scraping rules to make agents cheaper and more reliable. His friend's shopping agent used a browser for everything, hit CAPTCHAs, and wasted tokens. Rebuilding it, he replaces discovery's browser and fixed retailer list with Oxylabs' fast search API (under 2,000 tokens, under 700 milliseconds), and the decision stage with a scraper API that returns markdown, fails loudly, runs hundreds of parallel requests, and bills only for successful results. Checkout needs a browser, so Playwright MCP connects to Oxylabs' headless browser with stealth, residential proxy, and geolocation. Cost matters: use a browser only when necessary, validate content—HTTP 200 does not mean valid.

  1. 0:00The idea
  2. 2:32Scraping principles
  3. 3:59Agent setup
  4. 5:09Discovery
  5. 6:50Fast Search API
  6. 8:25Decision stage
  7. 10:29Scraper API
  8. 12:14Checkout
  9. 13:50Lessons

Powered by PodHood

Transcript

The idea0:00

Giedrius Šteimantas0:14

What a beautiful voice. Allright, thank you for coming. Um, today I'm going to talk a lot about missing layer of agentic AI and explain a little bit about how web scraping infrastructure can actually help you. But first, let me talk, uh, a little bit about my friend's idea.

So my friend had this idea, uh, he built this AI chatbot that, you know, chatted with people about their style, and it was supposed to help them pick out new items, uh, as s you know, some sort of a personal shopper.

And once those items were picked out, you know, this, uh, this this chatbot would, uh, produce prompts that a shopping agent would then take and attempt to find them online and purchase them for, uh, you know, for for the customers.

Um, this idea, you know, is not new and, uh, it could be applicable to many scenarios, but my friend was kind of, you know, uh, he was, uh, he was good at building agents, uh, but, uh, he ran into different problems and asked me for advice.

And when he ran it, he he would u-usually, you know, instead of, you know, product pages or whatever, he would get things like that. It's, uh, you know, he would get CAPTCHA. You know, and, uh, you know, of course, you know, he was, uh, he was doing it very, very quickly.

So he vibe coded the whole thing while having a, you know, a thought about, you know, infrastructure and underlying layers and how it should work at i at all. Uh, he was using a browser automation framework for everything, and it was slow, expensive, and unreliable.

So at the end, he made a product that, uh, uh, that does not work and is expensive to run. So he asked me for help. And, you know, I was a little bit reluctant at first because, uh, you know, I don't like giving out professional advice, you know, for free.

But, uh, I took a look at it and, uh, you know, I got a little curious, I have to be honest. I noticed that he was missing something. Um, he was missing a layer, an infrastructural layer that would allow this agent to operate freely on the open web.

Scraping principles2:32

Giedrius Šteimantas2:32

My name is Giedrius. I I work for Oxylabs, uh, where in the past 10 years, we've helped, you know, companies that trained large language models, uh, get their data, and now we use this infrastructure to help AI agents to access, uh, web on scale and at low cost.

And, uh, before we go into this agent and see how we can build it, I wanted to talk a little bit about the scraping industry and how we operate. And, uh, the principles that we operate on can be summed up by one, uh, sentence.

You know, cost matters. And the first principle is use a browser when you absolutely have to, validate content, HTTP response 200 does not mean that we are good to go. Lighter content is preferred. Websites are full of JavaScript, CSS, HTML, and there's a lot of bytes that do not deliver any value whatsoever.

And today, I will demonstrate how these principles are also applicable when building agents that interact with the web. So coming back to my friend's agent,right? Let's, uh, let's take a look and see how, uh, we could do a better job and, uh, making this agent run more reliably.

So here's how my friend set it all up, you know. So four different stages: discovery, the agent is was supposed to find product pages on websites where these items can be bought; then a decision stage,right, and, uh, where, uh, an agent can decide, uh, what products to buy based on, you know, uh, the c the content of these pages.

Agent setup3:59

Giedrius Šteimantas4:21

So the agent has to visit them, verify that the the stock is there, the price isright, the the the description fits, uh, you know, the prompt. And once that decision is made, user is given with a choice, you know, whether to go ahead with the purchase or, you know, reject it altogether.

The problem was that sometimes and, of course, we go to executionright away, then execution is making the purchase. But the problem was that sometimes it worked and sometimes it did not. That was a little problematic. So let's dissect it step by step and see how we could build this differently while improving performance and reducing the cost dramatically by using the s-same principles from the scraping industry.

Discovery5:09

Giedrius Šteimantas5:10

So the first stage, discovery. So my friend, uh, you know, he chose to go with a predefined list of websites, major retailers, uh, and query their search pages in order to find these products. He used a browser automation tool.

For that, it kind of worked, but, you know, it did have challenges. So their browser automation tool lacked what we call stealth. So they could so they would get CAPTCHAs and sometimes fail to access the sites altogether. This would break down the flow.

So a retry mechanism would have to be put in place, making the whole process very long, uh, you know, costly, uh, and sometimes these sites would not be, uh, accessed at all. And also, you know, as a result, it also became very difficult to predict the final cost per transaction.

The list of websites that my friend was checking was also deterministic. So selection of items would only be limited with the few choices he put in. Websites themselves were heavy on JavaScript, making the whole process very slow and costly.

And finally, even if it worked, items ended up being unavailable at checkout because in the discovery phase, the he was not able to use any geolocation capabilities, and a lot of e-commerce websites are, uh, you know, uh, they take your user's location into account when displaying stock, options, sizes, and so on.

So now, we solve these problems at Oxylabs every day. So when scraping, you always want the results to appear on the first try and to not use browser unless absolutely necessary. However, for this specific s uh, discovery phase, you also want to use to allow your agent to search the web.

Fast Search API6:50

Giedrius Šteimantas7:05

Doing so with a browser is very cumbersome. That is why I chose to use a product that we built, especially for agents, uh, fast search API. It returns a compact JSON, which is less than 2,000 tokens per response, has fast response times, less than 700 milliseconds on average, and it's, uh, has a a high success rate at a predictable low price.

And most importantly, it gives your a-agent access to the mo to, you know, to many popular search engines that all of these websites have been inxed and indexed already a long time ago. So in the discovery phase, instead of a predefined list and a browser, we give agent a tool to search the web, fast search API.

Agent fo-formulates fan-out queries and selects the relevant URLs from search results. Since the responses are quite small and there's no need for complicated models, we can have the agent run quite quickly in this stage.

Um, yeah. So, so now the agent has searched the web and selected some relevant URLs. It is time for those for for the agent to visit those pages to see what they're all about in order to confirm price, stock level, description, and product details, and so on.

Decision stage8:25

Giedrius Šteimantas8:25

With this, we can go to into decision phase. This is where agent selects the items we will purchase. For this, my friend also used the browser. He ran many browsers on parallel so it could, uh, you know, so the whole process could happen faster, and that is not a bad thing.

He managed to get some results. However, many of the results would end up like this.

And the result, the agent would be left with very few choices, with the majority of popular retailers being left out. It's a good thing he did well with observability, so he actually noticed when it happened. But what we see when working with these types of customers is that they often fail to detect the failure.

They end up checking only the content size and HTTP response code and then feeding this large HTML to an LLM. Now, and a large language model, of course, can distinguish between valid e-shop content and a CAPTCHA, but we need to spend tokens in order to do that.

And when we attempt to open 10 websites, but only 3 return valid content,

but feed all of the 10 to the to the model, it is a problem. It means that we waste 70% of the tokens. And that is a little crazy in my in my opinion.

So I noticed this problem as well. Uh, my initial hunch was compression, was to compress the output. But then I thought, wait, the problem is not the compression. The problem is that the content is not valid. We need to make sure that the content is valid be-be-before even attempting any compression.

This will lead to more options for the agent to choose from and fewer wasted tokens. And then I remembered rule number one of scraping. Use a browser when you absolutely need it. Otherwise, look for other solutions. So I I tried to rebuild this stage without a browser, and I, uh, only by using Oxylabs' web scraper API.

Scraper API10:29

Giedrius Šteimantas10:37

And this gave me many benefits. Uh, but firstly, only valid content was returned. In case of CAPTCHAs or other blocks, the request would fail with an explicit error message, so I know not to include it when sending to a large language model.

But the success rates are quite high, uh, even for protected websites, so that wasn't that much of a you know, much of a problem. So no browser was needed. And, uh, everything is a lightweight REST API. I can run hundreds of requests in parallel and receive content at the same time.

Also, the API supports markdown, so no need to sub uh, submit raw HTML, uh, to LLMs. If a website is dynamic, it runs a full browser under the hood to render the content correctly. And finally, it supports geolocation options, so I can localize my results and get relevant content.

The best part, customers only pay for successful results. So I'll actually, yeah. That's, uh, that's what's, uh, that's what's that's what's the best thing about it. No cure, no pay. If if the scraper fails, there's no cost, and it fails loudly.

So now we have all of the information to make a decision. We present a decision to the user, and the user makes the final call. Once it's affirmative, we move to the last stage of the workflow, the purchase.

So I remember what I said a couple of times about browsers. This time but this time is different. You this time, you absolutely need to use a browser. We need to process inputs, and the content is highly dynamic.

Checkout12:14

Giedrius Šteimantas12:25

Now, this time, my implementation, my friend's implementation, does not differ much. We both use Playwright MCP with a browser and a large language model.

The main problem my friend faced, however, just like in, uh, in the previous stages, uh, using browser, was access. Just like in the beginning, as he was using the browser, he was getting CAPTCHAed into oblivion, making it impossible to automate the flow.

Well, the fix was quite easy. I just connected Oxylabs' headless browser. Since it supports Playwright MCP, it's just a drop-in replacement. With this replacement, I hardened this agent with years of scraping experience and got proper stealth done at the browser source code level, a residential proxy attached to it out of the box, and most importantly, in this in this case, a geolocation capability.

So my results are localized the same way as in the verification stage. So if we run it, we actually have a a a a browser that that access the content and can actually automate the flow by, you know, selecting theright size from the prompt, add it to cart, and complete the purchase.

And boom, we have an agent that commands a powerful infrastructure hardened by years of web scraping experience. Not only does it open the up the web, but also saves the time on implementation and token cost. And if I can leave you with a few lessons we learned today, was that, you know, when building agents, use the same principles from the scraping industry.

Lessons13:50

Giedrius Šteimantas14:21

Use the browser when you absolutely need to. You have to validate content before feeding it to the, uh, large language models. And most importantly, fill the missing layer with the proper infrastructure so you can focus on building stuff.

But remember, cost matters. Thank you very much.