(I wasn't involved in building kitesurf, but I am informed that they intend to open source and upstream their patches)
[edit: for others reading who don't usually nerd out on browser automation protocols: webdriver bidi is the new-ish w3c cross-browser standard inspired by CDP - the main magic was the upgrade to websockets and also to standardize the capture of network-level traffic. there are still feature gaps between CDP and BiDi (in spec and implementation), but long term, i believe we should bet on web standards, not proprietary protocols controlled by one company.
(disclosure: i started the selenium and appium projects.)]
Is it a good idea to already build something on top of Blitz?
Just curious on your thoughts about how Webkit was architected then, I guess it's not a modular system where you can separate out things like "Localstorage" support?
These two feel like they are opposing teams, I don't think they are colluding today, but how long will that last, this seems very suspicious I say that as a long time cloudflare user, I welcome making the platform agent friendly and adding agent specific deployment cloud stuff like Cloudflare OS is something I can live with as well.
But this is going a bit too far, what's next AI bot net to scrape content from sites protected by Cloudflare? I don't want to sound entitled but man do we deserve better.
> Run headless Chrome on Cloudflare's global network for browser automation, web scraping, testing, and content generation.
Does Cloudflare the CDN allow these browser instances to bypass their own anti-bot mechanisms? Or will Cloudflare the CDN block them the same as if someone was running scraping bots from a different provider?
Will Kitesurf in Cloudflare workers get special bypass privileges to content protected by Cloudflare the CDN?
We also have a documented UA and sign our requests with Web Bot Auth: https://developers.cloudflare.com/browser-run/reference/auto...
My wife really dislikes building up the shopping cart for our weekly grocery delivery, so I built an agent... thing with earendil's npm libs. It takes the menu my wife has decided on, confers with her about the ingredients (if it hasn't seen a recipe before), and then uses Chrome's devtools protocol to head to Walmart and add everything to the shopping cart.
It works fairly well and uses the local models I have running on my Mac Studio.
I also tell my agents to remove annoyances from websites I browse, rearrange the content so that it's easier for me to view. For example when somebody publishes a table where they compare their newly released AI model to others I tell my agent to highlight highest result for each benchmark in every table on the page. I could do it myself with a bit of JS but why bother if agent can write it for me. I added a functionality to my agentic browser that lets the agent make userscripts for me that I can trigger with a push of a button.
I also ask agents whether the specific information is on the page that I'm currently browsing in language I don't understand (or just among the clutter).
Once I asked agent to put more than a dozen items into a cart for me (which names I pasted) because the ecommerce site didn't have convenient way of doing that.
So basically Grease Monkey on steroids + TD;DR;whaat?
Local Qwen3.6 is smart to do all that but I have option to switch to remote stronger models.
I didn’t “use an agent to find a receipt” in the sense that I purpose built one. I just asked my existing agent that I talk to on telegram by photographing the thing I wanted to know if we could return and while I changed the baby it chugged along and by the time we were ready to go it could tell me whether we did buy it at Costco and when so I know if I can return it.
- Apple Appstore Connect (gazillions of forms of metadata to release an app) - AWS - DigitalOcean - Google Play Store
Whenever I dread logging in because I know the simple sounding task requires me to click through countless menus I use an agent browser. With confirmations of course. However, while the agent clicks through these (oftentimes dog slow) UIs I can do other things. Once it requires permission, I read, decide and act.
tldr; to workaround the lack (or shortcomings) of public m2m APIs in web apps
From what I understand, Kitesurf eventually went in its own direction and isn’t simply Obscura running on Workers. Still, knowing that Obscura helped get the original experiment started means a lot.
I recently added native rendering to Obscura. It can now take screenshots, stream screencasts, and generate PDFs without Chromium. The repository has also passed 21,000 stars.
For context, I’m 16 and have mostly been building this with my friend, so I’m still figuring out what the project should become and how to keep developing it sustainably.
Happy to answer any questions about Obscura or the rendering work.
time for another approach to run your agent's web searches through this (or by mocking browser signature), with potential cf bypass built-in!
We have been keeping a close eye on BiDi as well.
It's a web data tool, but as something not used for browsing, by definition this is not a browser.
a welcome addition although it'd be very easy for websites to fingerprint and block
You can block today, Kitesurf doesn't try to hide.
https://developers.cloudflare.com/browser-run/reference/auto... https://developers.cloudflare.com/browser-run/faq/
Results from the last hour, verbatim:
- lemmy.world /api/v3/site: registration_mode RequireApplication, captcha_enabled true, require_email_verification true, and an application question that explicitly rejects temporary email. Three independent walls on one signup. - lemmy.today and lemy.lol /api/v3/user/register: {"error":"captcha_incorrect"}. The captcha ships as base64 PNG plus WAV, so it is a wall for anything without a decoder, headless browser or not. - bsky.social com.atproto.server.createAccount: {"error":"InvalidPhoneVerification"}. - Publishing, by contrast: api.telegra.ph and write.as both take an unauthenticated POST and hand back a public URL.
A browser in a V8 isolate does not help with any of the failures above, because the gate is a CAPTCHA, an SMS, or a card on file, and an isolate has none of those. The same is true on the payments side: an agent can hold an address and receive, but every write path in that ecosystem is a signature over a payload, so if something else custodies your key the machine-payments world is read-only to you.
The missing primitive for agents is not a browser. It is a portable identity and a spendable balance that are not borrowed from a human's phone and credit card.
Full map of what was reachable and what was not: https://write.as/ih3l0kd78lpb1