37 points•STRiDEX•about 4 hours ago•46 comments•

46 comments

dvtabout 4 hours ago
I built an AI "web harness" running on a sandboxed Chromium (using a custom side-loaded plugin that talks over websockets to a "driver") to basically do anything a normal user could do in a browser. It totally bypasses any and all bot measures and only gets the ones you yourself would get as well (and passes those successfully, e.g. Cloudflare checkbox or those annoying OCR puzzles).

Not sure if I should release it, but I'm sure more people are catching onto the power of agentic browsing.

bayindirhabout 3 hours ago
Thanks for letting us know that we need a new layer of detection systems.

Also it’s great(!) to see that we’re going from “but ethics” to “I got mine, who cares”.

Humans are interesting creatures.

Edit: Please before assuming that I'm assuming things, this is an observation I'm making over time. It's possible that I'm in a bubble, but it's not a sample size of 1 (i.e. The comment I replied only).

afro88about 3 hours ago
I'm becoming more and more convinced that a big source of outrage on the internet is caused by people assuming that all other people are a homogenous blob.

It's not that "we" are going from one thing to another. It's that these are two different people, with different ethical boundaries.

pcthrowawayabout 3 hours ago
There is no detection method that will prevent AI from accessing systems without also blocking humans. The only thing we can do at this point is throttling.
Cakez0rabout 3 hours ago
You're framing this as if people are deliberately making decisions that they believe are unethical. The reality is that people have different ethical frameworks. For example, I believe that there is no ethical distinction between whether a web request originates from a browser or from an LLM on my behalf.
Cakez0rabout 3 hours ago
I think it's inevitable that eventually one of the big AI companies makes something like this. It's such an obvious consumer win. "I am not a bot" checkboxes are a string and peg tethering the elephant.
onion2kabout 3 hours ago
(using a custom side-loaded plugin that talks over websockets to a "driver")

Is that necessary? You could start Chromium with an open debug port and use Chrome Devtools Protocol to send commands.

dvtabout 3 hours ago
It is, because `--remote-debugging-port` is detected via the root DOM object (and that can't be changed unless you want to recompile Chromium), so for example, if you try doing a Google search with debugging enabled, you'll get blocked (usually just by being served a blank page).
pprotasabout 3 hours ago
Camoufox bypasses most blocks with no problems https://camoufox.com/
kurisufagabout 3 hours ago
the 4get dev did the same thing a little while ago for his metasearch engine: https://git.lolcat.ca/lolcat/4play
dvtabout 3 hours ago
Yep, it's super similar to this! I think his is a bit overengineered, but tbh I haven't built Firefox plugins in forever so that might be the way you have to do stuff on FF. All you need is the websocket plugin to have "full access" to all visited webpages, and then your driver just hooks into the DOM like normal (with the extra ws layer—and even the websocket layer can be simplified away I think).
wietherabout 3 hours ago
As someone having to fight Meta's bots everyday to keep websites accessible to actual customers, I'm not surprised to read that it can be seen as something positive on the other side of the fence.

But I'm wondering: at what cost?

Cakez0rabout 2 hours ago
The end game will be that either your site is fully open to bots and humans alike, or your site is open to humans only(×) and requires Airport security style identity verification.

(×) and their AI delegates

STRiDEXabout 3 hours ago
I think for my own side projects i would require the user to login if they were making requests from those ip addresses or block.
kevmo314about 3 hours ago
Muse ran into a captcha and asked me if I wanted it to solve it.

So of course I clicked yes and it dutifully convinced the site that it was not a bot.

xnickbabout 3 hours ago
For 2(3?) decades we've been training the robots to tell traffic lights from fire hydrants. It's finally paying off.
Gareth321about 3 hours ago
I think this is the end of sites where it is expected that only humans may interact with them. It's been cat and mouse for a while, and some places like Reddit sell API access, but these agents are for all intents and purposes, humans interacting with the site. They're going to need to figure out new business models.
mejutocoabout 3 hours ago
Or the captchas will evolve.
gunalxabout 3 hours ago
Im guessing more of the internet will be login walled from now on.
arjunchintabout 3 hours ago
their static ip's were initially good and didn't get flagged, but now most sites are recognizing their ip ranges and blocking.

Muse's utility has significantly dropped with the blockages.

To become truly useful again they will need to use residential proxies, but I can't see them use those due to the risks and reputational damage.

wraptileabout 2 hours ago
Every time a new tool launches there's a good window where it can act as a scraping proxy. Back in the 2010s I used Google Translate for years to scrape hard targets like LinkedIn but these windows are much shorter these days as scraping is so much bigger.

One thing with Muse though is that you can scrape Meta's own sites which are currently all going under login walls and restricting discovery/search entirely.

STRiDEXabout 2 hours ago
similarly, gemini via google ai studio will happily run workloads over youtube that would be very annoying to run at scale. Especially if you needed to download the video.
Gareth321about 3 hours ago
Most residential proxies are already far more blocked and rate limited than any Meta IP. The internet is becoming a very weird place, where individual and "trusted" personal IPs are becoming a kind of commodity. Some sites are already scoring IPs based on usage activity - like a credit score. It's only a matter of time until this data is collated and commoditised. AI analysis is turning this up to 11.
iamacyborgabout 3 hours ago
That doesn’t track with reality as far as I can tell.
simoncionabout 3 hours ago
> Some sites are already scoring IPs based on usage activity - like a credit score. It's only a matter of time until this data is collated and commoditised.

Spamhaus is nearly thirty years old and the notion of electronic distribution of IP and domain "reputation" lists is at least that old.

I'll bet my hat that the Internet "advertising" [0] industry has been calculating and determining the reputation of individual households (if not individual users) for at least a decade.

[0] The scare quotes are because its primary purpose these days is for dragnet private-sector surveillance.

sejjeabout 3 hours ago
they can just use the user ip. i think grok already does this.
koolalaabout 3 hours ago
How can it do this? Wouldn't you see a hundred fetch requests in your browser network tab?
arjunchintabout 2 hours ago
bruh have you heard of CSP?

Read the full thread on Hacker News →

Related stories