data
289 stories and discussions about data, aggregated from every source we track.
A sample of 5,000 alleged agents seen by 404 Media includes names, addresses, phone numbers, and details on FBI employees' spouses.
Exfiltrate LLM weights and data through GET requests
One file, entire app. Share like a document, open like an app.
Companies behind major negligent data leaks can now face fines of up to 10 percent of annual revenue under revised privacy rules.
In 51,129 comments from six knife subreddits, the 5% most brand-heavy accounts wrote 11.3% of the brand mentions in buying threads, against 7.9% expected by chance. For one brand it is 31% against 8%. The full Reddit…
Police didn't care that Linsey Isaacs' car was the wrong color, and wasn't damaged. They still arrested her because of Flock data.
Data centre sizes and costs are doubling about every 12 to 16 months
It sends encrypted signals with no satellites or relay towers required
The vast majority of data centers in Europe keep information about their environmental impact secret, according to a study by Lighthouse Report in collaboration with Trouw and other European media. In the Netherlands,…
Using Facet's run-time reflection to check the layout of structs against their definitions in WebGPU shaders
The Data Protection Commission (DPC) has today announced its final decision following an Inquiry into Google Ireland Limited (“Google”).
Such strikes aim to disrupt "people's ability to stay connected, study, work", says Ukraine's president.
There is an API. It will tell you things about your writing that the dashboard will not. I have been...
The agreement between LinkedIn, ProAPIs and joint business operator Netswift also requires the firms to stop selling and transferring the data, no longer access LinkedIn through fake accounts and delete the data that…
SB 923, signed September 27, 2026, extends the CCPA right to delete to data obtained from third parties and requires online-only businesses to offer a web form. It takes effect January 1, 2027.
Two commands against the same data. Run them yourself: $ curl -sS --get...
Google's experimental orbital data center will have four TPUs and only run for 15 minutes at a time.
Is data duplication always bad? How to deal with the additional invariants in the relational table design?
Food prices today for onions, potatoes, tomatoes, lettuce, avocados and strawberries: USDA wholesale and shipping point quotes with price charts, change against a year earlier, grocery ad prices by region and daily…
Jev is a useful zero-shot classifier, but its probabilities can't be calibrated for your data. Calibration depends on your data distribution, which Jev never sees, so treat its outputs as scores and recalibrate them on…
XGo is a programming language that reads like plain English. But it's also incredibly powerful — it lets you leverage assets from C/C++, Go, Python, and JavaScript/TypeScript, creating a unifie...
What happens when we collect too much data?
Prepared for the Kaggle Benchmarking Challenge. What I Benchmarked I build agents for...
Building an Interactive South America Dashboard with Streamlit and World Bank Data Open...
Benchmarks & Tips for Big Data, Hadoop, AWS, Google Cloud, PostgreSQL, Spark, Python & More...
Waves, wind, water and tides from NOAA's network of buoys, coastal stations, tide gauges and estuaries, from the Bering Sea to Samoa, charted and analysed.
Learn about cluster mirroring, the embedded cross-cluster replication in Kafka, and its benefits for disaster recovery and cluster migration.
The single requirement of all data pipelines is that they cannot lose data. Data can usually be delayed or re-ordered–but never dropped.
To all the self-proclaimed ‘experts’ here who spent the weekend defending Google, calling my forensic proof ‘AI-generated slop’, and excusing a fake delete button as ‘standard industry practice’: you might want to read…
I measured PostgreSQL 19 beta's online checksum conversion against the old offline tool: roughly 1.3s of hard downtime the old way versus 2.8s of background work and 799MB of extra WAL the new way, with zero drop in concurrent write throughput.
Wavelets have become a hot topic in mathematics in general and computer graphics in particular. They are a useful tool for multiresolution analysis of a data set and for data compression. There are several popular…
We built a multi-tenant team task board with no backend. The browser loads static files, signs in...
Million-dollar fines won’t end billionaires’ data center pollution, neighbors fear.
You know it's cold when you have to heat the air used to cool your data center.
Matrix, the open protocol for secure decentralised communications
Last September, we introduced an optional new way to automatically keep your Signal data safely and securely backed up so that important messages and memories can be easily recovered if something unexpected happens to…
National Center for Education Statistics, which collects critical data to guide policy on learning, was decimated after Doge canceled crucial contracts
Two providers changed the rules, and my app went back to the hangar Showtime on the...
I work as a data architect in San Francisco & Dr. Jones mentioned you might be able to help me before I get too deep into the design of a new system. My default database choice is to just use Postgres. I have questions…
A Deep Investigation I've been sitting on this project for a while, I haven't shared it...
Datastory is the no-code platform for data storytellers. Create interactive charts, websites, and reports using Studio, CMS, AI, and 2000+ open datasets.
New study finds cars send driver data to Big Tech companies, including locations, VINs, and other identifiers.
Europe's AI data center buildout is creating a power problem that cannot be solved by adding more...
The problem If you've built anything with LLMs, you've hit this wall: your data lives in...
<p>From the Abstract:</p> <blockquote> <p>Localizing type errors is challenging in languages with global type inference, as the type checker must make assumptions about what the programmer intended to do. We introduce Nate, a data-driven approach to error localization based on supervised learning. Nate analyzes a large corpus of training data — pairs of ill-typed programs and their “fixed” versions — to automatically learn a model of where the error is most likely to be found. Given a new ill-typed program, Nate executes the model to generate a list of potential blame assignments ranked by likelihood. We evaluate Nate by comparing its precision to the state of the art on a set of over 5,000 ill-typed OCaml programs drawn from two instances of an introductory programming course. We show that when the top-ranked blame assignment is considered, Nate’s data-driven model is able to correctly predict the exact sub-expression that should be changed 72% of the time, 28 points higher than OCaml and 16 points higher than the state-of-the-art SHErrLoc tool. Furthermore, Nate’s accuracy surpasses 85% when we consider the top two locations and reaches 91% if we consider the top three.</p> </blockquote>
ERP software handles money, stock, and approvals, so small bugs get expensive fast. A stock count...
Get a free Jylus API key, send your own logs, telemetry or JSON, and inspect compact source-backed evidence.