Amazon left relational. Google came back. Both were right.
Contents· The table you came for
- 1The table you came for
- 2“NoSQL can’t do transactions”
- 3“SQL can’t scale horizontally”
- 4“NoSQL is schemaless”
- 5One fact, seen from two ends
- 6Why Amazon left
- 7Why Google came back
- 8The two lists do not disagree
- 9“Pick two of three”
- 10Neither side is the safe one
- 11“Cheaper at scale” has no general answer
- 12You are picking a point on a line
- 13Every claim, traced
Most of what you believe about SQL and NoSQL came from blog posts written around 2012. Those posts were right at the time. They are not right now. And the two companies everyone quotes as proof both wrote down what actually happened.
The table you came for
Every one of these comparisons ends at a scorecard. Here is the one you probably carry in your head. Three of its four rows are wrong. The fourth is the interesting one.
| Claim | SQL | NoSQL | Status today |
|---|---|---|---|
| yes | no | wrong since 2018 | |
| no | yes | wrong — three counterexamples | |
| no | yes | wrong both ways round | |
| Joins at the database | yes | no | true — and it is the whole point |
“NoSQL can’t do transactions”
NoSQL can't do transactions.
Wrong — MongoDB has had them since 2018
MongoDB shipped multi-document ACID transactions in version 4.0, in 2018. Sharded clusters got them in 4.2. A transaction can span operations, collections, databases and shards. It all commits, or none of it does.
That is easy to over-read, so here is MongoDB telling you not to:
In most cases, multi-document transactions incur a greater performance cost over single document writes, and the availability of distributed transactions should not be a replacement for effective schema design.
“SQL can’t scale horizontally”
SQL can't scale horizontally.
Wrong — three of them, on three different engines
Not one exception. Three, and they sit on three different relational engines.
| System | Built on | What it proves |
|---|---|---|
| Spanner | Google-built | SQL, horizontal scale, and a guarantee stricter than serializable — held “across an entire database (even across multiple Cloud regions) without blocking writes.” |
| Citus | PostgreSQL | “an extension (not a fork) to Postgres — when you use Citus, you are also using Postgres.” You shard without leaving the database you know. |
| Vitess | MySQL | “served all YouTube database traffic for over five years.” |
“NoSQL is schemaless”
NoSQL is schemaless.
Wrong at both ends at the same time
MongoDB never called itself schemaless. The docs say “flexible schema model”. Ask it to enforce one and it will:
By default, MongoDB rejects any insert or update operation that would produce an invalid document — you can also set validation to warn instead.
Now look the other way. PostgreSQL stores JSON in jsonb, indexes it with GIN, and queries it with a path language. That is a document store with SQL bolted to the front. One side never made the claim. The other side went and became it.
One fact, seen from two ends
So what is different? Take one order. Store it both ways. Then do the two things you will actually do to it.
| What you do | Normalised (3 tables) | Embedded (1 document) |
|---|---|---|
| Read order 91 | 3 tables, 2 joins | 1 read |
| Rename the Widget | 1 row | every document that copied it |
Why Amazon left
Amazon never wrote “relational databases don’t scale.” They named three specific problems. The specifics are the useful part:
Amazon learned that providing applications with direct access to traditional enterprise database instances led to scaling bottlenecks such as connection management, interference between concurrent workloads, and operational problems with tasks such as schema upgrades.
DynamoDB started as one job: “a highly scalable, available, and durable key-value database for shopping cart data.” It ended up peaking at 89.2 million requests per second over the 66-hour Prime Day 2021 event, at single-digit millisecond latency.
Why Google came back
Ten years earlier, on a different system, Google went the other way — away from a wide-column store, back toward SQL and transactions. Same reason: specific complaints.
we have also consistently received complaints from users that Bigtable can be difficult to use for some kinds of applications: those that have complex, evolving schemas, or those that want strong consistency in the presence of wide-area replication.
And on the old argument that transactions are too expensive to offer, they took the other side:
We believe it is better to have application programmers deal with performance problems due to overuse of transactions as bottlenecks arise, rather than always coding around the lack of transactions.
The two lists do not disagree
Put them side by side and the argument goes away. These are not two answers to one question. They are two teams describing two different workloads. That is your decision criterion, and you did not have to invent it — they published it.
Which six do you have?
tick what is true of your system today
Amazon's three
Google's three
Nothing ticked yet
Tick the ones that are true for you. This points, it does not decide. It knows six published problems and nothing about your team, your latency budget, or what you already run.
“Pick two of three”
Pick two of three.
The man who wrote the theorem disowned it
The “2 of 3” formulation was always misleading because it tended to oversimplify the tensions among properties.
He did not take the theorem back. He took the slogan back. The tradeoff is real; the triangle is what misleads. Three corrections matter far more than the shape:
| What the slogan says | What Brewer actually wrote in 2012 |
|---|---|
| You choose once, for the whole system | You choose per operation. Different subsystems, different operations, even different rows can choose differently. |
| A partition is a broken cable | A partition is a deadline: “Failing to achieve consistency within the time bound implies a partition and thus a choice between C and A for this operation.” |
| Give up C or A forever | Partitions are rare. Giving either one up while everything is healthy buys you nothing. |
Neither side is the safe one
Kyle Kingsbury’s Jepsen reports are the closest thing this field has to a referee. He has caught both sides.
PostgreSQL’s “serializable” isolation level isn’t serializable: it allows G2-item during normal operation.
That is a real quote. Stopping there would be dishonest. The same report says it was a bug — it had been there since serializable snapshot isolation shipped in 2011, across versions 9.5 to 13 — and PostgreSQL found the cause, patched it, and shipped the fix within weeks. That is a database being tested and repaired, not a database being unsafe.
Jepsen also tested MongoDB 4.2.6. It failed snapshot isolation even at the strongest read and write concerns: read skew, cyclic information flow, duplicate writes. That result is from 2020, and it keeps its version number for the same reason. Quoting a five-year-old finding as if it were today is the identical mistake, aimed the other way.
“Cheaper at scale” has no general answer
I researched this one and then cut it. The two things are not measured in the same unit. DynamoDB charges for “each read or write request consumed.” Aurora charges “per instance-hour consumed,” with storage and I/O billed separately. One meter counts requests. The other counts time.
Worse, each vendor sells a mode that flips the comparison — DynamoDB has provisioned capacity by the hour, Aurora has I/O-Optimized. However you set up the sum, the other side has a setting that undoes it. So the claim is not in this piece. One narrower thing is true and useful: a strongly consistent read costs more capacity than an eventually consistent one. The CAP tradeoff is not only in your architecture. It is on your bill.
You are picking a point on a line
PostgreSQL got binary JSON with indexing in 2014, then sharding through Citus. MongoDB got ACID transactions in 2018, then schema validation. The gap that 2012 advice described has been closing since the day it was written. That is why advice from that year now misses in both directions at once.
So stop asking which one is better. Ask what your app does most. Ask whether your reads look like your writes. Ask whether your schema has stopped moving yet.
Amazon and Google asked those questions, got different answers, and both showed their working. That is the part worth copying.
Every claim, traced
The chips beside each section are not decoration. Not all evidence is worth the same, and a piece that treats a conference paper and a vendor’s pricing page as the same thing has already lost the argument.
- peer reviewedAmazon DynamoDB: A Scalable, Predictably Performant, and Fully Managed NoSQL Database Service ↗Elhemali et al. — USENIX ATC 2022
- peer reviewedSpanner: Google's Globally-Distributed Database ↗Corbett et al. — OSDI 2012
- peer reviewedCAP Twelve Years Later: How the “Rules” Have Changed ↗Eric Brewer — InfoQ, 2012. The theorem's author on his own slogan.
- independent testingJepsen: PostgreSQL 12.3 ↗2020
- independent testingJepsen: MongoDB 4.2.6 ↗2020
- project docsMongoDB Manual — Transactions ↗
- project docsMongoDB Manual — Schema Validation ↗
- project docsPostgreSQL — JSON Types (jsonb, GIN, SQL/JSON path) ↗
- project docsSpanner — TrueTime and external consistency ↗
- project docsVitess — What is Vitess ↗
- project docsCitus — What is Citus ↗
- vendor figuresAWS — DynamoDB on-demand pricing ↗
- vendor figuresAWS — Aurora pricing ↗