Two BM25 optimizations and benchmark configuration changes inspired by PlanetScale's TIN benchmarks make ParadeDB's text search faster without changing its document identifiers.

43 points•craigkerstiens•about 3 hours ago•8 comments•

8 comments

jamesgresqlabout 2 hours ago
ParadeDB'er here: We are really hoping this can turn into a back and forth, making both of our products better.
esafakabout 2 hours ago
If you modify it for use in a commercial backend (not a hosted db service), does the AGPL-3.0 license mean you have to share the source of project it is used in or only of the forked repo? On a related note, have you changed your stance on accepting issues and pull requests like some orgs?
philippemnoelabout 2 hours ago
As for the AGPL question, our primary goal with this license is to protect ourselves from hyperscalers and avoid fragmenting the community. You can read more about it here: https://news.ycombinator.com/item?id=41227172
philippemnoelabout 2 hours ago
We have many community contributors (150 and counting) who submit PRs and many users who open issues. We love them!

We're well aware that our external contributors use AI (we do too!), and we fiercely guard our product against AI slop. We ask that external contributors understand the changes they are making, and we frequently close PRs unmerged if they look like slop.

peterpanheadabout 2 hours ago
ParadeDB kicks the TIN can should be the title.
tschellenbachabout 3 hours ago
Definitely looking at using them, you do wonder if such narrow focus is defensible as a business? Probably not.
philippemnoelabout 2 hours ago
ParadeDB developer here. BM25 FTS is becoming available on more cloud platforms, but ParadeDB is much more than that now. We'll have blog posts over the next few months detailing our work on aggregates, JOINs, filters, etc.
meeritaabout 1 hour ago
Massive respect for ParadeDB for seeing a benchmark where they were 8x slower, taking it as an engineering challenge instead of making excuses, profiling the hell out of it, fixing the actual bottlenecks, and publishing everything they learned along the way.

This is what healthy engineering competition looks like.

Read the full thread on Hacker News →

Related stories