Add Scry to ChatGPT or Claude and ask in your own words. It reads the forums, the papers, the filings and the markets, and comes back with the answer and the links.

61 points•Xyra•13 days ago•26 comments•
Meet Scry, a 500 TB NVMe internet index in ClickHouse that you can run ~arbitrary readonly SQL and some of Datalog over, and I handle the problem of resource-contention with congestion-based micro-auction pricing. When there's capacity, the service is free for non-commercial use.

---

Hello. It's 2026, we're training simulated fruit fly brains to play Beat Saber, do we still have to be stuck with internet (re)search as fn: natural language -> black box we can't do anything about -> ranked_list/summary?

There is a long history of people trying to do very fancy things that end up being done in relational databases and a little SQL. There is a gravity to them, a bitter lesson, just like scaling of generalized ml training methods. I mean many, many information products can be built off essentially giant real-time OLAP databases and frontier LLMs writing brilliant SQL+Datalog+vector+Jev etc. queries.

Google Search, Tavily, Exa essentially have the problem of mapping your agents' context you are willing to provide, to a tiny subset of their index. You pay a fixed cost to an extremely hard problem that has a distribution of hardness, which means YOU eat the downsides when they are running out of budgeted compute to help you out.

Their algorithms are opaque to the caller, there's really not much user control, and there's not a serious opportunity to communally improve search recipes, like the lexical+Jev recipes you trust to select bleeding edge AI builders.

Furthermore, search companies aren't even pursuing text-to-SQL anymore (several have talked to me)... they made up their minds during the traumatic 2024 text-to-sql days. They were just too early.

I hope you enjoy. I'm intent on scaling this paradigm on differentiated hardware over much more data, so any compelling use cases or queries I could show off, would be much appreciated!

26 comments

neilellis12 days ago
Please have a 'readable version' option so I don't have to exhaust myself parsing the sites layout. I get that it's unique but most of us just want to work out what you're offering in 5-10 seconds of our time.
JLO6412 days ago
I strongly second this, although I must admit it loaded surprisingly fast for me as I'm on a mobile hotspot in the back of a car.
hmartin12 days ago
HN: This site looks like all the other slop, awful to read.

Also HN: This site is doesn't look like other sites, awful to read.

ai-inquisitor12 days ago
I'd bet that no one, not even the site's [human] creators, has ever read that homepage end to end. At best, it might have been handed over to a swarm of reviewer agents.
ashkankiani12 days ago
Cool idea. The pricing model reminds me of my time working in algorithmic trading, haha. I'll try this out for some queries I wanted to run.

I suppose the scraping you're doing is a huge part of your value proposition, but I would like to gently nudge you in the direction of making the datasets available via p2p (e.g. a torrent) like how Wikipedia distributes its snapshots in the spirit of democratizing access to data that is becoming increasingly walled off. Also, I think another potential benefit that kind of bulk sharing would have is relieving the congestion from those doing the equivalent operation to extract data via the querying interface.

Xyra12 days ago
let me know if it was faster and more compositionally expressive than you expected
mrbluecoat12 days ago
Scry, keyword action that allows a player to look at a specific number of cards from the top of their library and then arrange those cards in any order, placing any number of them on the bottom of the library and the rest on top.
codexon12 days ago
How did you scrape reddit comments? Doesn't this require expensive licensing from reddit? How do you handle comments that were deleted by users?
saturatedfat12 days ago
pushshift baby
codexon11 days ago
How hasn't reddit sued them offline yet?
NoahZuniga9 days ago
I was actually able to find something with this that I would have never been able to find with traditional search! This is great.

Read the full thread on Hacker News →

Related stories