Why DuckDB 2.0 is faster(motherduck.com) |
Why DuckDB 2.0 is faster(motherduck.com) |
> One setting drives this,...
> The cost is now about the rows you actually touch, not rounds times table size.
Etc.
I get the brain scramblies [1] from trying to parse this writing style at work so I hate to see it elsewhere. Apologies if I'm wrong. But if I'm not then OP don't use an LLM to write for you. It's hazardous to your reader's health [2].
[1]: https://www.youtube.com/watch?v=ipUJq-odt5Q
[2]: https://discourse.haskell.org/t/how-to-keep-enjoying-program...
> Storing a lake as thousands of 1 MB Parquet files is a bad practice anyway, and 2.0 does not rescue it.
The "does not rescue it". No human would write like that.
> I'll explain what that means on a table you already know.
No I don't already know that table.
Also
> and claims 40x on graph reachability
Is really hard to parse.
The section on recursive CTEs wasn't well written and didn't explain how the optimisation was done. This article explains how the recursive CTEs were improved https://duckdb.org/2026/08/25/how-duckdb-runs-recursive-ctes...
This is what I don’t understand. Supposedly LLMs are trained on human text. Why do they come up with such unrealistic prose? Is it intentional because the companies want the tells to be obvious?
On the contrary. It is unnatural for a native speaker, which I don't think the author is. For someone that speaks English as a second language, it is not uncommon to use expressions literally translated from their first language, which may be understandable but weird for native speakers.
I'm starting to read it as a sign of low LLM effort not just low human effort. It seems most common when one few-sentence prompt leads it to generate 4+ paragraphs (and the longer the output, the worse the odds). Prompting to dig into each resulting paragraph one by one, to make them readable, makes it do higher-effort deep dives.
Just because some guy on some forum said that doesn't make it true. That's not how you establish facts regarding health claims.
Not that i question that reading mostly ai slop for long enough makes you feel dead inside.
I get the brain scramblies [1] from trying to parse this writing style at work so I hate to see it elsewhere
This sounds to me a lot like mass hysteria, people reading other people's behaviour online and reproducing it unconsciously.More and more really important and useful information will arrive like this for us to consume. There's no way around. So this is a disservice for newcomers that could come and go unscathed but instead is crippled by these kinds of comments that brings nothing of substance to the table and has the potential to make them hate something they otherwise wouldn't even notice.
Also, from https://news.ycombinator.com/newsguidelines.html:
Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something.
Please don't complain about tangential annoyances—e.g. article or website formats, name collisions, or back-button breakage. They're too common to be interesting.This sort of writing decreases readability. People say "just have your own LLM rewrite it" like you don't lose value when you go from prompt -> slop you didn't review enough to clean this crap up in -> someone else's prompt to change the style -> finally someone reads it.
Do you know what happens a double-digit percentage of the time when I ask Claude to rewrite some shit that it gives me like this? It says things like: "I overstated this, I rechecked and actually..." or "this claim doesn't hold up, actually [this other thing is true]..."
So it's a sign that the claims in the post likely weren't vetted very hard.
So if you aren't proofreading I'm gonna be skeptical. And saying "deal with it" doesn't rescue it. Does it?
And then there's the reflexive "you must just be an ideological hater." No. I'm someone who uses the tools in a domain where quality matters enough that I have to dig into the quality of the tool output and spot the tells for when it's low output, so that I can deliver shit that works reliably and consistently.
Most of the DB engines out there still seem to use a "n-threads" style parallelism with exchange operations and poor async I/O management.
DuckDB is improving on this front, but in some sense is catching up to R&D (and implementation!) that is now decades old.
A "rhetorical challenge" I like to give software developers working on systems like this is the following: If I gave you a computer with 1,024 cores and matching network and storage bandwidth -- but with significant latency -- could you keep a system like this 100% utilised with one query?
The answer for almost all software is "no".
For example, SQL Server tops out at 64 hardware threads for any one query: https://learn.microsoft.com/en-us/sql/database-engine/config...
GPU codes are starting to get there, but CPU codes are way behind on this frontier of computer science.
It's not just databases! Can you (de)compress a file in parallel? Verify its hash in parallel? Upload/download from storage with CPU and I/O task parallelism? Can you overlap all of these operation so nothing is ever waiting on anything else it doesn't have to?
This matters! I ran some tests with bioinformatics codes and found that most got stuck in tar pits. Many could not scale to modern SSDs with millions of IOPS or modern networking with hundreds of gigabits of throughput, no matter how many CPU cores were thrown at them.
PS: AMD's Zen 6 era EPYC 9006 processors will have 512 cores and 1,024 threads per two-socket system, so this is not hypothetical: https://www.amd.com/en/products/processors/server/epyc/9006-...
So it's a sign that the claims in the post likely weren't vetted very hard.
Our time is better used submitting something useful instead of debating meaningless guesswork in the comments.Also if you think LLMisms so bad it's spam, you also don't get a free pass. From the guidelines:
If a story is spam or off-topic, flag it. Don't feed egregious comments by replying; flag them instead. If you flag, please don't also comment that you did.