The physical reality behind the AI-compute story.
Public notes on where the economics of AI infrastructure get decided. That is rarely the headline compute number. It is the precision the hardware is built for, the watts and heat it moves, and the clock it depreciates on. The aim is to make the reasoning legible enough that the next question gets sharper.
A build or an investment is priced on headline FLOPS, while power, cooling, utilization, or the depreciation clock are treated as details rather than the decision.
The precision the workload needs is drifting away from the hardware roadmap; the real constraint is watts and heat, not compute; the asset is depreciating faster than the model assumes.
Public architecture specs, system datasheets, depreciation disclosures, reliability literature, and public power-market data. Ranges, not point forecasts.
Which single number would move the decision: utilization, power price, depreciation life, or aging failure. Another headline throughput figure will not.
People who need the AI-compute story to survive contact with watts, heat, depreciation, and utilization.
Operators, founders, engineers, and commercially minded technical people around AI infrastructure, high-performance computing, and the power and storage that sit under it. Long-lived capital, fast-moving hardware, and the uncomfortable truth that "compute is abundant" is true for some workloads and false for others.
The asset is physical and financial at once
At what utilization, power price, and depreciation life does the hardware pay? And who is answering the physical half of that question?
Precision is not free
Which workloads are safely low-precision, and which structurally need FP64? Where does the hardware roadmap help, and where does it stop helping?
Watts and heat set the limit
Power and cooling, not the card's sticker price, decide where a build can run and whether it runs at all. They also set how fast it ages.
The GPU roadmap is racing down the precision ladder, stranding the science that needs full double precision.
As I read the public architecture specifications, each recent data-center GPU generation packs more of the low-precision arithmetic that model training and inference reward: FP16 and BF16, then FP8, now FP4. FP64, full double precision, has not scaled at the same pace. On the newest AI-first parts it occupies a small corner of the chip. This is a directional reading of public specs; check the current primary sources (NVIDIA architecture whitepapers, vendor datasheets) before any decision.
Down the ladder, for good reasons
Neural networks tolerate low precision well, so it is rational to spend transistors where the AI workload pays back. For most of the AI market this is the right call.
FP64 is physics, not habit
CFD, structural FEA, climate and weather, molecular dynamics, and computational materials lean on FP64 for conditioning. Stiff systems and long integrations accumulate rounding error that lower precision cannot absorb. The result is a wrong answer that still looks plausible.
Two curves drifting apart
Per watt, accelerators get faster at the math AI needs and slower at the math a jet engine or a new alloy needs. "Compute is abundant" is true for AI. It is much less true for double-precision science.
Open question worth holding: which FP64 workloads are non-negotiable, and which survive mixed precision and iterative refinement? The honest answer separates a real infrastructure gap from an engineering habit. My own numerical-simulation work uses FP64 solvers over real geometry, so I have a stake in the answer.
The AI-capex story is priced on compute; the risk lives in three quieter numbers.
Watts
A single dense AI server, eight accelerators plus the rest of the box, pulls the power of a small building at full tilt. Public system datasheets put top-end boxes around the ten-kilowatt mark. Power and cooling, not the card's price, set where a build can run.
Heat
Every watt becomes heat that has to leave. Sustained high power density is what ages the hardware: thermal cycling, solder fatigue, ordinary wear. Run it hot and busy and it does not only cost more to cool. It dies sooner.
Time
A frontier accelerator can be the best in the world and still depreciate on a short clock, because the next generation resets the frontier. Accounting life, physical life, and market-value life are three different numbers. The gap between "obsolete for training" and "still useful for something" is worth more than most models assume.
Directional, public-source framing (system datasheets for power; public depreciation disclosures and reliability literature for aging). FLOPS purchased is the wrong measure. The asset pays at some combination of utilization, power price, depreciation clock, and aging failure. That combination is half physical and half financial, and the two halves are usually answered by different people.
What this note is, and what it is not.
Claims are dated and public
Hardware roadmaps, power figures, and depreciation conventions age quickly. A useful note keeps the source line close to the conclusion and never launders a private number into a public one.
Ranges stay ranges
Where the evidence is directional, the language stays directional. False precision about watts, prices, or depreciation is worse than an honest band.
Public note, not advice
Nothing here is a recommendation on any specific build, asset, or transaction. A real decision still needs primary specs, measured power, and current source checks.
Better evidence moves the note
New silicon, new datasheets, and disproved assumptions are supposed to change the weight of what is written here. Corrections are welcome.
Questions, corrections, or public-source pointers are welcome.
Especially useful: better primary sources on FP64 throughput and system power, cleaner terminology, or cases where a physical reality was priced into an infrastructure decision, or missed.
Die physische Realität hinter der KI-Compute-Geschichte.
Öffentliche Notizen dazu, wo die Ökonomie von KI-Infrastruktur entschieden wird. Selten in der Schlagzeilen-Rechenleistung. Entscheidend sind die Präzision, für die die Hardware gebaut ist, die Watt und die Wärme, die sie bewegt, und die Uhr, auf der sie abschreibt. Ziel ist es, die Argumentation lesbar genug zu machen, dass die nächste Frage schärfer wird.
Ein Aufbau oder eine Investition wird auf Schlagzeilen-FLOPS bepreist, während Strom, Kühlung, Auslastung oder die Abschreibungsuhr als Detail statt als Entscheidung behandelt werden.
Die benötigte Präzision driftet weg von der Hardware-Roadmap; die echte Grenze sind Watt und Wärme, nicht Rechenleistung; der Vermögenswert schreibt schneller ab als das Modell annimmt.
Öffentliche Architektur-Spezifikationen, System-Datenblätter, Abschreibungsangaben, Zuverlässigkeitsliteratur und öffentliche Strommarktdaten. Bandbreiten, keine Punktprognosen.
Welche einzelne Zahl würde die Entscheidung bewegen: Auslastung, Strompreis, Abschreibungsdauer oder Alterungsausfall. Eine weitere Schlagzeilen-Durchsatzzahl tut es nicht.
Menschen, die die KI-Compute-Geschichte den Kontakt mit Watt, Wärme, Abschreibung und Auslastung überstehen lassen müssen.
Betreiber, Gründer, Ingenieurinnen und kommerziell denkende technische Menschen rund um KI-Infrastruktur, Hochleistungsrechnen und die Strom- und Speicherbasis darunter. Langlebiges Kapital, schnelllebige Hardware und die unbequeme Wahrheit, dass „Rechenleistung ist im Überfluss vorhanden" für manche Lasten stimmt und für andere falsch ist.
Der Vermögenswert ist zugleich physisch und finanziell
Bei welcher Auslastung, welchem Strompreis und welcher Abschreibungsdauer zahlt sich die Hardware? Und wer beantwortet die physische Hälfte dieser Frage?
Präzision ist nicht gratis
Welche Lasten sind sicher niedrigpräzise, und welche brauchen strukturell FP64? Wo hilft die Hardware-Roadmap, und wo hört sie auf zu helfen?
Watt und Wärme setzen die Grenze
Strom und Kühlung, nicht der Listenpreis der Karte, entscheiden, wo ein Aufbau laufen kann und ob überhaupt. Sie bestimmen auch, wie schnell er altert.
Die GPU-Roadmap rast die Präzisionsleiter hinunter und strandet dabei die Wissenschaft, die volle doppelte Präzision braucht.
Wie ich die öffentlichen Architektur-Spezifikationen lese, packt jede jüngere Rechenzentrums-GPU-Generation mehr der niedrigpräzisen Arithmetik ein, die Modelltraining und -inferenz belohnen: FP16 und BF16, dann FP8, jetzt FP4. FP64, die volle doppelte Präzision, hat nicht im gleichen Tempo skaliert. Auf den neuesten KI-zuerst-Chips belegt sie einen kleinen Teil des Chips. Dies ist eine richtungsweisende Lesart öffentlicher Spezifikationen; vor jeder Entscheidung die aktuellen Primärquellen prüfen (NVIDIA-Architektur-Whitepapers, Hersteller-Datenblätter).
Die Leiter hinunter, aus guten Gründen
Neuronale Netze vertragen niedrige Präzision gut, daher ist es rational, Transistoren dort auszugeben, wo die KI-Last es zurückzahlt. Für den Grossteil des KI-Markts ist das die richtige Wahl.
FP64 ist Physik, keine Gewohnheit
CFD, strukturelle FEA, Klima und Wetter, Molekulardynamik und computergestützte Werkstoffkunde stützen sich für die Konditionierung auf FP64. Steife Systeme und lange Integrationen häufen Rundungsfehler an, die niedrigere Präzision nicht auffangen kann. Das Ergebnis ist eine falsche Antwort, die weiterhin plausibel aussieht.
Zwei Kurven driften auseinander
Pro Watt werden Beschleuniger schneller bei der Mathematik, die die KI braucht, und langsamer bei der Mathematik, die ein Triebwerk oder eine neue Legierung braucht. „Rechenleistung im Überfluss" stimmt für KI. Für doppelt-präzise Wissenschaft stimmt es weit weniger.
Offene Frage, die es zu halten lohnt: welche FP64-Lasten sind nicht verhandelbar, und welche überleben gemischte Präzision und iterative Verfeinerung? Die ehrliche Antwort trennt eine echte Infrastruktur-Lücke von einer technischen Gewohnheit. Meine eigene numerische Simulationsarbeit nutzt FP64-Löser über realer Geometrie, ich habe also ein Interesse an der Antwort.
Die KI-Investitionsgeschichte wird auf Rechenleistung bepreist; das Risiko lebt in drei leiseren Zahlen.
Watt
Ein einzelner dichter KI-Server, acht Beschleuniger plus der Rest der Kiste, zieht unter Volllast den Strom eines kleinen Gebäudes. Öffentliche System-Datenblätter setzen Spitzenkisten um die Zehn-Kilowatt-Marke. Strom und Kühlung, nicht der Kartenpreis, setzen, wo ein Aufbau laufen kann.
Wärme
Jedes Watt wird zu Wärme, die abgeführt werden muss. Anhaltend hohe Leistungsdichte altert die Hardware: Temperaturwechsel, Lotermüdung, gewöhnlicher Verschleiss. Heiss und ausgelastet betrieben kostet sie nicht nur mehr Kühlung. Sie stirbt früher.
Zeit
Ein Spitzenbeschleuniger kann der beste der Welt sein und dennoch auf einer kurzen Uhr abschreiben, weil die nächste Generation die Grenze neu setzt. Buchhalterische, physische und Marktwert-Lebensdauer sind drei verschiedene Zahlen. Die Lücke zwischen „veraltet fürs Training" und „noch für etwas nützlich" ist mehr wert, als die meisten Modelle annehmen.
Richtungsweisende, öffentlich belegte Einordnung (System-Datenblätter für Strom; öffentliche Abschreibungsangaben und Zuverlässigkeitsliteratur für Alterung). Gekaufte FLOPS sind das falsche Mass. Der Vermögenswert zahlt sich bei einer Kombination aus Auslastung, Strompreis, Abschreibungsuhr und Alterungsausfall. Diese Kombination ist halb physisch und halb finanziell, und die beiden Hälften beantworten meist verschiedene Menschen.
Was diese Notiz ist und was nicht.
Claims sind datiert und öffentlich
Hardware-Roadmaps, Stromzahlen und Abschreibungskonventionen altern schnell. Eine nützliche Notiz hält die Quellenlinie nah an der Schlussfolgerung und wäscht nie eine private Zahl in eine öffentliche.
Bandbreiten bleiben Bandbreiten
Wo die Evidenz richtungsweisend ist, bleibt die Sprache richtungsweisend. Falsche Präzision über Watt, Preise oder Abschreibung ist schlechter als eine ehrliche Bandbreite.
Öffentliche Notiz, keine Beratung
Nichts hier ist eine Empfehlung zu einem bestimmten Aufbau, Vermögenswert oder einer Transaktion. Eine echte Entscheidung braucht weiterhin Primär-Spezifikationen, gemessenen Strom und aktuelle Quellenprüfungen.
Bessere Evidenz bewegt die Notiz
Neues Silizium, neue Datenblätter und widerlegte Annahmen sollen das Gewicht des hier Geschriebenen verändern. Korrekturen sind willkommen.
Fragen, Korrekturen oder öffentliche Quellenhinweise sind willkommen.
Besonders nützlich: bessere Primärquellen zu FP64-Durchsatz und System-Strom, sauberere Terminologie oder Beispiele, bei denen eine physische Realität in eine Infrastrukturentscheidung eingepreist wurde, oder eben übersehen.