Empirical Laws of AI
Scaling Laws for Neural Language Models
How model size, data, and compute collapsed onto a small set of power laws—and reorganized the logic of training language models.
Open full explainerOne operator. One semantic engine. It turns a frontier-physics paper and an enterprise go-to-market brief into the same kind of source-anchored interactive instrument — then wires them into one live correlation matrix.
One engine · two ends of the spectrum
The instrument doesn’t care whether the source is a theoretical-physics preprint or a competitive positioning brief. It reads the source, finds the one load-bearing reversal, builds a thing you can manipulate, and refuses to invent what the source didn’t say. The register is identical at both ends.
Cool pole · the papers
Two of these explainers descend from Juan Maldacena’s holographic duality — the result that information on a boundary can encode an entire bulk. That lineage runs straight into kinematic space and the observer papers.
Warm pole · the campaigns
The identical discipline — thesis-first, source-anchored, epistemically bounded — turns a messy AI-security story into positioning, campaign architecture, and interactive proof surfaces a leadership team can actually test.
01
Paper, manuscript, or campaign brief in. Every consequential claim gets a verbatim locator before anything is built.
02
The one move that changes what counts as the object — place → interval, model size → scaling law, spam filter → agentic perimeter.
03
A coding-agent prompt specifies a Next.js page you can manipulate — deterministic state, inline SVG, no invented data.
04
The page emits its semantic anchor and joins the correlation matrix, so the set gets smarter with every node.
The operator behind the instrument
This is the working proof of an AI-native content operating model: one senior operator converting dense research and ambiguous market change into strategy, interactive artifacts, and public proof surfaces — at a velocity and surface area that used to require a full agency pod. The through-line from Maldacena to marketing isn’t a metaphor. It’s the same instrument, pointed at both ends.
Context Jamming / Explainers
Foundational papers and texts from AI and physics, rebuilt as interactive research instruments. Manipulate a core diagram from each source, then open the full explainer for the argument, equations, and source notes.
Empirical Laws of AI
How model size, data, and compute collapsed onto a small set of power laws—and reorganized the logic of training language models.
Open full explainerCompute-Optimal Scaling
Allocate a fixed compute budget between parameters and training tokens, inspect the IsoFLOP valley, and see why the 70B Chinchilla beat the 280B Gopher.
Open full explainer Kaplan 2020 companion →Euclidean Gravity
Why de Sitter’s state count rotates through i—and why the observer must enter the calculation before we decide what is being counted.
Open full explainerKinematic Space
How a tensor network became a map of every possible way to probe a universe.
Open full explainerConstitutional Classifiers++
How Anthropic made a stronger jailbreak defense approximately 40× cheaper by moving classification into the generation stream.
Open full explainerEntropic Gravity
How a local quantum-relative-entropy action between two metrics produces a dressed route back to Einstein gravity.
Open full explainerSpecial Relativity
How an operational definition of time resolved the conflict between light and relativity—and forced space-time.
Open full explainerUniversal Weight Subspaces
If training repeatedly converges into a narrow geometry, can gradient-aligned filtering and constraints stop the optimizer wasting the trip?
Open full explainerGenerative Pre-Training
The live GPT-1 paper explainer: pre-train a Transformer on unlabeled text, then adapt it with supervised fine-tuning—and inspect exactly which weights move.
Open full explainer Alec Radford profile companion →In-Context Learning
How scaling to 175B parameters steepens few-shot learning curves—competitive with fine-tuned SOTA on many tasks, with frozen weights and no gradient updates.
Open full explainer Oral History companion →Mechanistic Interpretability
How path expansion turns attention-only transformers into readable QK and OV circuits—and how two layers compose a previous-token head into the induction algorithm.
Open full explainer 2022 induction-heads follow-up →Mechanistic Interpretability
The ICL score jump, induction-head formation, a loss-curve bump, and a PCA pivot all happen in the same narrow training window — six lines of evidence for the mechanism behind most of what prompting does.
Open full explainerWeight-Space Geometry
A trained network as a collectible action figure: roughly sixteen shared directions per layer act as engineered joints, and 1,100+ models across architectures turn out to be different poses on the same skeleton.
Open full explainer Filtered-training companion →AI Alignment
How a short written constitution, self-critique, and AI preference labels replace human harm labels — producing an assistant that engages and explains instead of going silent.
Open full explainer Classifiers++ successor →Unsupervised Intelligence
How a single bounded-observer quantity—spectral epiplexity—resolves the noisy-TV and dark-room pathologies, turning noise into solitons, digit clusters, and stable exploration.
Open full explainerLong-Context Attention
How segment-level recurrence and relative positional encodings first gave pure self-attention long-range memory — 0.99 bpc, 18.3 PPL, and up to 1,800× faster evaluation.
Open full explainerLong-Context Memory
How checkpointing a recurrent memory turns the RNN/Transformer choice into a single tunable coordinate — and recovers gated softmax attention at the far end of the dial.
Open full explainer Transformer-XL predecessor →Research Transformation System
Next-generation skill. Turns any arXiv paper into a sleek Next.js interactive research dashboard (HTML prototype + full App Router implementation). Also handles SaaS enterprise marketing campaigns → interactive pages with ROI labs, persona switchers, and objection simulators. Epistemically bounded, source-anchored, production-ready.
Open full explainerSynthesis layer
Select two to four papers. The matrix distinguishes structural affinity from documented influence, revealing shared mechanisms, conceptual bridges, and the points where the comparison breaks down.
Method note. Affinity scores are editorial synthesis aids, not statistical measurements. Direct influence is shown only when supported by a citation or explicit source evidence.
| Paper | Scaling Laws2020 | GPT-3 Few-Shot2020 | Constitutional AI2022 | Classifiers++2026 |
|---|---|---|---|---|
| Scaling Laws2020 | Self Compress training behavior into empirical power laws. | |||
| GPT-3 Few-Shot2020 | Self Scale until tasks are learned in context, with frozen weights. | |||
| Constitutional AI2022 | Self Swap human harm labels for a short constitution plus AI feedback. | |||
| Classifiers++2026 | Self Move full-context judgment into an adaptive compute cascade. |
Scaling Laws × GPT-3 Few-Shot
Scaling laws established that loss falls predictably with model size, data, and compute; GPT-3 is the bet that riding those curves far enough would buy qualitatively new behavior — and few-shot in-context learning is what it bought.
Measure the trend at small scale, trust the regularity, and spend unprecedented compute on the extrapolation.
Documented influence
The GPT-3 paper explicitly discusses and cites Kaplan et al. (2020), and the arXiv author records overlap substantially (Kaplan, Brown, Amodei, and McCandlish appear on both). This is the most directly documented lineage in this matrix.
Scaling laws predict smooth loss curves; GPT-3's headline phenomenon — in-context learning — is exactly the kind of qualitative jump the loss curve does not directly predict.
A power law in cross-entropy is not a theory of few-shot learning; the laws constrain loss, not which abilities appear when.
Documented direction: the scaling-law results provided the quantitative case for training a 175B-parameter model at all.
Retrospective comparison only: GPT-3's few-shot results then became the canonical evidence that the scaling program pays off in capabilities, not just loss.
Selected-set synthesis
93 / 100 affinity
80.7 average affinity
Most recurrent curated tags in the selected pair records
Dates order the comparison; they do not establish influence.
Generated deterministically from the selected set’s curated tags.
From Kinematic Explainer to Production Dashboard
One reusable skill turns a dense paper or enterprise campaign record into a source-bounded implementation brief for an instant HTML prototype and a production-grade interactive Next.js experience.
This is the natural power-up of the Kinematic Space method: the same thesis-first, source-anchored discipline, expanded from a single Context Jamming explainer into a portable dashboard system for research, product, and go-to-market work.
A self-contained first pass using Tailwind CDN, inline SVGs, and deterministic local state for immediate browser review.
A Next.js 14+ App Router TypeScript implementation with server components, client islands, URL-shareable state, and KaTeX-ready equations.
Global controls drive linked visualizations, filterable evidence tables, stress-test labs, and epistemic boundaries expressed directly in the interface.
A new sub-skill for ROI/TCO calculators, persona labs, journey simulators, objection handlers, and testable attribution models.
What’s new in v2
Invocation examples
Use the attached arXiv paper and the research-to-interactive-dashboard skill to generate a complete copy-paste prompt for a sleek Next.js interactive research dashboard (HTML prototype first, then full production implementation).Use this SaaS campaign brief and the research-to-interactive-dashboard skill (campaign sub-skill) to generate an interactive Next.js marketing page with a live ROI simulator and persona lab.Download and unzip the package, then place the skill folder where your coding agent can read it.
Attach an arXiv paper, technical manuscript, or enterprise campaign brief with the source material that governs the build.
Invoke the skill, review its 18-part implementation prompt, then give that prompt to Codex, Cursor, or another coding agent.
Epistemic layering remains non-negotiable: established results, author interpretation, and dashboard or strategic extensions stay visibly distinct. Synthetic reconstructions are labeled, analogies expose their break points, and every consequential claim retains a source anchor.