CONTEXT JAMMING

Field notes from inside the context window.

Share

FounderFiles·N°061·Physics · Scaling · Reasoning

2008 —

Ethan Dyer, PhD — theoretical physicist and scaling-law theorist; formerly Google Blueshift, now at Anthropic
Fig. · The cartographerEthan Dyer. Photograph: Gabrielle Lurie/Quanta Magazine

Subject·Ethan Dyer, PhD·Theoretical physicist · Scaling-law theorist · Former Google Blueshift, now Anthropic

Ethan Dyer, PhD

Dyer took the physicist’s toolkit of Feynman diagrams, large-N limits and phase diagrams and pointed it at neural networks — then helped build the benchmark and the math model that showed what scale could and couldn’t buy.

He began as a holographer counting black-hole microstates in toy universes. At Google he drew Feynman diagrams for wide networks, found the learning rate at which training “catapults,” co-wrote with Jared Kaplan the paper that explains where scaling laws come from, helped organize the 204-task stress test that made “emergence” a headline, and helped build the language model that learned to do math. Then, quietly, he turned up at Anthropic — the same company as his co-author.

TRAINED
Columbia · MIT (PhD, 2014) · Stanford (postdoc)
AT
Anthropic (per Aspen, 2025–26)
FILE
N°061
§ 01 · The Physicist

From a boundary term to a boundary theory

His first paper was about boundaries. As a Columbia student in 2008, Dyer co-wrote with Kurt Hinterbichler Boundary terms, variational principles, and higher derivative modified gravity (Phys. Rev. D, January 2009). The paper asks what you must add at the edge of spacetime to make a theory of gravity’s equations well-posed. A companion paper worked out the same question for the DGP “π-Lagrangian” and the galileons. It is still one of his most-cited physics papers, at roughly 270 Google Scholar citations.

He went to MIT for his PhD (2009–2014) and joined the Center for Theoretical Physics. His MIT-era work sits in the AdS/CFT world: monopole operators in three-dimensional conformal field theories with Márk Mezei and Silviu Pufu, and super-Rényi entropies and Wilson loops for N = 4 super-Yang–Mills and their gravity duals with Michael Crossley and Julian Sonner (JHEP, 2014). The title and advisor of his dissertation do not appear in any public thesis record we could find, so this file does not print them.

He then moved to Stanford’s Institute for Theoretical Physics (2014–2017) and went further into three-dimensional gravity and two-dimensional CFT: An Extremal N = 2 Superconformal Field Theory and Universal Bounds on Charged States in 2d CFT and 3d Gravity with Nathan Benjamin, Liam Fitzpatrick and Shamit Kachru; 2D CFT partition functions at late times with Guy Gur-Ari; Spinning Geodesic Witten Diagrams with Daniel Freedman and James Sully; Constraints on Flavored 2d CFT Partition Functions with Fitzpatrick and Yuan Xin, which lists both Stanford and Johns Hopkins; and The most irrational rational theories(2019). He gave a Princeton seminar titled “Small black holes in near extremal gravity,” on counting black-hole microstates in the conjectured duality between pure 3D gravity and “extremal” 2D CFTs.

This is the same subject area as Kaplan’s Aspects of Holography: lower-dimensional boundary theories that encode higher-dimensional gravity. Dyer and Kaplan were not just two physicists who both drifted into machine learning. They were two holographers.

§ 02 · The Crossing

Stanford, Hopkins, and a workshop on physics for machine learning

The crossover can be dated. Dyer sat on the scientific organizing committee of a “Theoretical Physics for Machine Learning” workshop, listed as “Stanford University & Johns Hopkins.” His fellow organizers were Adam Brown, Paul Ginsparg, Guy Gur-Ari and Jaehoon Lee. By 2018 he was at Google. His first ML paper, Gradient Descent Happens in a Tiny Subspace (December 2018, with Gur-Ari and Dan Roberts), showed that during training the gradient quickly settles into a small subspace spanned by the top Hessian eigenvectors. It is a spectral argument of the kind a physicist would make, and it has about 375 citations.

The Johns Hopkins affiliation matters for the Kaplan story. Kaplan joined the Hopkins Department of Physics and Astronomy in 2012–13 and stayed on its faculty after joining OpenAI in 2019. Dyer carried a Hopkins affiliation on papers from about 2017 to 2019, so the two overlapped there. Neither has described that overlap publicly in anything we found, and there is no co-authored physics paper between them. Liam Fitzpatrick, a frequent co-author of both, links the two holography networks.

Why physicists made this crossing at all — and who else did — is the subject of The Data Desert: Physics, AI, and the Diaspora, where Kaplan is the defector and Dyer turns out to be a second one, with a co-author.

§ 03 · The Diagrams

Feynman rules for infinitely wide networks

The paper that made Dyer’s name in ML was Asymptotics of Wide Networks from Feynman Diagrams, written with Gur-Ari (arXiv:1909.11304; ICLR 2020). It brings a field theorist’s bookkeeping to neural networks. Treat the width as a large parameter, like N in a large-N gauge theory. Expand correlation functions of the network’s outputs and derivatives in 1/width. Organize the terms with diagrams and read off how each quantity scales before computing anything. It has about 159 citations.

Related papers followed. Asymptotics of Wide Convolutional Neural Networks, with Anders Andreassen, extended the program to CNNs. The large learning rate phase of deep learning: the catapult mechanism (arXiv:2003.02218, March 2020), with Aitor Lewkowycz, Yasaman Bahri, Jascha Sohl-Dickstein and Gur-Ari, identified a regime above the standard “lazy” learning-rate limit. There, loss first rises, curvature collapses, and the network “catapults” into a flatter region and generalizes better. That paper has about 366 citations.

This is Kaplan’s instinct — networks are physical systems, and physical systems have laws — done at the level of dynamics instead of empirical fits. Kaplan measured the curves. Dyer wrote down the perturbation theory.

“How we understand these sharp transitions is a great research question.”
Dyer, Quanta Magazine, March 2023
§ 04 · The Four Regimes

Explaining Neural Scaling Laws

February 12, 2021. Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee and Utkarsh Sharma posted Explaining Neural Scaling Laws (arXiv:2102.06701). A revised version appeared in PNAS121(27) in 2024. This is the link that holds this file together. It is the only verified paper Dyer and Kaplan share, and it arrived a year after Kaplan’s OpenAI Scaling Laws paper, as its theoretical counterpart.

The claim, from the abstract: variance-limited and resolution-limited scaling behavior for both dataset and model size, for a total of four scaling regimes.

Variance-limitedscaling follows from the existence of a well-behaved infinite-data or infinite-width limit. These are corrections to a smooth limit — the same large-N logic as the Feynman-diagram paper. Resolution-limitedscaling is explained by positing that models are effectively resolving a smooth data manifold; the exponent is set by the manifold’s intrinsic dimension, and higher-dimensional data means slower power laws. In the large-width limit the same exponents come from the spectrum of certain kernels, and the authors present evidence that large-width and large-dataset resolution-limited exponents are related by a duality. Model size and dataset size act as mirror images.

Kaplan’s 2020 paper said that loss follows power laws. Bahri, Dyer, Kaplan, Lee and Sharma said why, and what sets the exponent.

The Kaplan file calls the shape of the power-law claim the most consequential physics result in machine learning history. If so, this paper is its microscopic explanation, and Dyer is one of its five authors — with roughly 720 citations. Play with the curves in the interactive Scaling Laws explainer.

Three months after the preprint, Dyer took the argument to a room of theoretical physicists. His May 2021 talk, “The role of scale in deep neural networks”, is billed with an abstract about understanding how performance improves with scale — and about telling apart the problems scale alone can solve from the ones where new ideas are needed. That second clause is the one worth keeping in view: the theorist of why the curves exist was already asking where they stop.

§ 05 · The Imitation Game

The benchmark that made “emergence” a headline

In 2020, per Quanta Magazine, Dyer and others at Google Research predicted that large language models would have transformative effects, and asked the research community to contribute examples of difficult and diverse tasks. The result was BIG-bench, Beyond the Imitation Game(arXiv:2206.04615; TMLR 2023). The paper’s abstract counts 204 tasks, contributed by 450 authors across 132 institutions.

The finding that went viral: around 5% of BIG-bench tasks see models achieve sudden score breakthroughs with increasing scale, though that behavior can depend sharply on the metric used to probe performance. The paper’s contributions section says BIG-bench was managed and organized by Guy Gur-Ari, Jascha Sohl-Dickstein, Noah Fiedel and Ethan Dyer, and that BIG-bench Lite was developed by Dyer.

This is where the Dyer file meets the Kaplan file’s “Emergence as Mirage.” The same Quanta piece reports that when Dyer’s team posed the emoji-movie task as multiple choice, the accuracy improvement was less of a sudden jump and more of a gradual increase — the metric-artifact explanation that the 2023–24 “mirage” papers later made formal. BIG-bench’s data supported both the emergence story and its refutation.

Dyer’s own framing was open-ended: “sharp transitions” as a research question, not a proof of magic.

“Despite trying to expect surprises, I’m surprised at the things these models can do.”
Dyer, Quanta Magazine, March 16 2023
§ 06 · Minerva

Teaching a language model to show its work

June 2022. Solving Quantitative Reasoning Problems with Language Models (arXiv:2206.14858; NeurIPS 2022) introduced Minerva. Dyer is one of fourteen authors, and the Google Research blog post was bylined “Ethan Dyer and Guy Gur-Ari, Research Scientists, Google Research, Blueshift Team.” It describes a model that solves mathematical and scientific questions using step-by-step reasoning, producing numerical calculations and symbolic manipulation without relying on external tools such as a calculator.

Related 2022 papers from the same team: Exploring Length Generalization in Large Language Models (NeurIPS 2022 oral), which documented how badly models extrapolate to longer problems than they were trained on; Block-Recurrent Transformers (about 220 citations); and Effect of scale on catastrophic forgetting (ICLR 2022, about 313 citations). Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models (2023) followed on the same line of work.

Dyer presented this line on November 15, 2022, at a Stanford Applied Physics/Physics Colloquium (video) titled “Lessons from scale for large language models and quantitative reasoning.” The abstract notes that some measures of progress show remarkably robust power-law improvement over many orders of magnitude, while other capabilities remain difficult to extrapolate. That sentence joins ENSL and BIG-bench in a single line.

§ 07 · Blueshift to Gemini to Anthropic

The quiet move

Blueshift became the reasoning group behind Gemini. Behnam Neyshabur, who co-led it from 2022 to 2024, writes that Blueshift was responsible for the reasoning capabilities of the first version of Gemini, and that it went on to develop a Gemini 1.5 math-specialized model and improve reasoning in Gemini 1.5 Pro and Flash and Gemma 2. Dyer is a credited author on the Gemini 1.0, 1.5 and 2.5 technical reports. Author lists that large, in the thousands, do not show individual roles, so his specific Gemini contribution is not established. As late as the September 2024 Michelangelo long-context evaluation paper, he appears as a contributor listed under Google DeepMind.

Then the affiliation changes. Neyshabur left Google for Anthropic in late 2024 to co-lead a Discovery team aimed at building an AI scientist and engineer; he has since left to co-found Mirendil. The Aspen Center for Physics lists “Ethan Dyer, Anthropic”as an organizer of the January 11–16, 2026 “Theoretical Physics for Artificial Intelligence” meeting, alongside Adam Brown, Dmitry Krotov and Eva Silverstein. Anthropic science posts from 2026 thank “Ethan Dyer” for feedback — “Long-running Claude for scientific computing” (March 23, 2026) and “Paving the way for AI agents in biology” — and his Hugging Face profile lists Anthropic as his organization.

What the record does not give us: his start date, title, or team at Anthropic. That he works on the Discovery or science side is an inference from the acknowledgments and from Neyshabur’s path, not a confirmed fact.

Nor does anything tie him to Gemini 4 Argon. Google’s September 30, 2026 announcement names no individual contributors, and Dyer was listed at Anthropic by then. Argon ships the reasoning line his old team started; the person who helped start it works for the rival.

§ 08 · The Convergence

Two holographers, one company

The paths did not diverge into rival labs, as the obvious frame assumes. They diverged for about five years — Kaplan building a lab, Dyer building reasoning systems inside Google — and then convergedat Anthropic. The side-by-side is below. Dyer’s move also fits a documented 2024–26 flow of Blueshift and Gemini reasoning talent toward Anthropic, including his long-time collaborator Neyshabur; as it applies to Dyer, that pattern is an inference.

Physics subjectKAPLAN · AdS/CFT, conformal bootstrap, amplitudesDYER · AdS3/CFT2, modular bootstrap, extremal gravity, black-hole microstates
Academic baseKAPLAN · SLAC & Stanford; Johns Hopkins faculty from 2012–13DYER · Stanford SITP 2014–17; Johns Hopkins affiliation ~2017–19
Entry to MLKAPLAN · OpenAI, 2019DYER · Google, 2018
Signature moveKAPLAN · Empirical power laws — Scaling Laws, Jan 2020DYER · Perturbative theory of networks; the why of the power laws, 2021
On emergenceKAPLAN · Smooth-loss view, later vindicated by the mirage papersDYER · Organized BIG-bench; the benchmark recorded both the breakthroughs and the metric artifact
2021–25KAPLAN · Co-founder and CSO of Anthropic; Responsible Scaling OfficerDYER · Blueshift → Gemini reasoning; Minerva; Gemini report credits
2026KAPLAN · AnthropicDYER · Anthropic (per Aspen; start date unverified)
Public profileKAPLAN · High — Guardian, TIME100 AIDYER · Low — Quanta quotes, conference roles, two recorded talks

The two holographers who co-wrote the theory of why scaling laws exist now work at the same company. That, not rivalry, is the shape of this story.

Timeline
  • 2008–09Columbia — first papers with Kurt Hinterbichler on boundary terms in modified gravity and galileons.
  • 2009–14MIT, Center for Theoretical Physics — PhD work in the AdS/CFT world. Thesis title and advisor not publicly verified.
  • 2014–17Stanford Institute for Theoretical Physics — postdoc; extremal and 3D gravity, 2D CFT partition functions.
  • ~2017–19Johns Hopkins — affiliation on papers, in the same department as Jared Kaplan.
  • 2018Google Research — joins the Blueshift team. Gradient Descent Happens in a Tiny Subspace (Dec).
  • Sep 2019Asymptotics of Wide Networks from Feynman Diagrams (ICLR 2020).
  • Mar 2020The catapult mechanism — a learning-rate “phase” above the lazy limit.
  • 2020BIG-bench effort launched at Google Research.
  • Feb 12 2021Explaining Neural Scaling Laws posted with Bahri, Kaplan, Lee and Sharma.
  • May 17 2021Gives the talk “The role of scale in deep neural networks” for the University of Chicago theoretical physics community (uploaded May 21).
  • Jun 2022Minerva; BIG-bench paper posted; length generalization and Block-Recurrent Transformers at NeurIPS.
  • Nov 15 2022Presents “Lessons from scale for large language models and quantitative reasoning” at the Stanford Applied Physics/Physics Colloquium (recording uploaded Nov 16).
  • Mar 2023Quanta Magazine quotes Dyer on emergence and “sharp transitions.”
  • Apr 2024Explaining Neural Scaling Laws published in PNAS.
  • Sep 2024Michelangelo long-context evaluation — contributor, listed under Google DeepMind.
  • Late 2024Behnam Neyshabur, Blueshift co-lead, leaves Google for Anthropic.
  • Jul 2025Credited on the Gemini 2.5 technical report.
  • By ~Sep 2025Listed as “Ethan Dyer, Anthropic” among organizers of the Aspen Center for Physics winter 2026 meeting.
  • Jan 11–16 2026Aspen — Theoretical Physics for Artificial Intelligence.
  • Mar 2026Thanked in Anthropic’s “Long-running Claude for scientific computing.”
  • Sep 30 2026Google announces Gemini 4 Argon. The announcement names no individual contributors.
The Index
4
Scaling regimes in ENSL — variance- or resolution-limited, data or model
1
Verified paper Dyer and Kaplan share
~5%
BIG-bench tasks showing sudden “breakthrough” jumps
204
Tasks in BIG-bench, from 450 authors at 132 institutions
2009
Year of Kaplan’s PhD and Dyer’s first paper
~34,000
Google Scholar citations, as captured
2018
Year Dyer joined Google; Kaplan joined OpenAI in 2019
0
Named individual contributors in Google’s Gemini 4 Argon announcement
Dossier

Education.Columbia University (undergraduate physics research, 2008–09). MIT, PhD in physics, Center for Theoretical Physics (2009–2014; dissertation details not yet verified). Postdoc at the Stanford Institute for Theoretical Physics (2014–2017).

Affiliations.Johns Hopkins Physics & Astronomy (affiliation around 2017–19). Google Research, Blueshift Team (2018 to about 2024), later listed under Google DeepMind. Anthropic (by about 2025; role and start date unconfirmed).

Collaborators worth naming. Guy Gur-Ari (co-founder of Augment), Aitor Lewkowycz, Yasaman Bahri, Jaehoon Lee, Jared Kaplan, Utkarsh Sharma, Jascha Sohl-Dickstein, Behnam Neyshabur, Vinay Ramasesh, Liam Fitzpatrick, and Adam Brown, with whom he co-organized both physics-for-ML meetings.

Range. Beyond physics and ML, a genomics paper (Tanigawa, Dyer, Bejerano, PLoS Computational Biology, 2022) and a glass-physics paper, Linking dynamical heterogeneity to static amorphous order.

Public voice. Thin on the record: two Quanta quotes (2023), the Stanford colloquium (Nov 2022, embedded at the top of this page), a 2021 University of Chicago theoretical-physics talk, the Minerva blog byline, and conference organizing roles.

Further reading
“The Unpredictable Abilities Emerging From Large AI Models”
Quanta Magazine · Mar 16, 2023 · where Dyer talks BIG-bench and emergence
“The Data Desert: Physics, AI, and the Diaspora”
Context Jamming · the physicists who left for AI — Kaplan, and Dyer after him
Career Shape
π-shaped — two deep spikes bridged by a general layer

π-Bridge

Carries the prior of a first field into a second and finds the governing law that was invisible to native practitioners; pays in delayed gratification.

Credential Path
Doctoral
Abstraction
Top Down
Exit Horizon
Deferred
Moat Instinct
Theoretical Insight
Capital Posture
None
Role-Model Reference Class
  • Large-N and holographic theorists
  • Jared Kaplan (co-author)
  • The physics-of-deep-learning tradition
Founder Context · JSON

A small reasoning persona distilled from this file. Inject it into a chat or deep-research context to assess a business problem the way Dyer would.

Reason as a field theorist assessing a business problem. First ask which limit the system is near: is performance capped by noise around a well-understood limit (variance-limited) or by how finely you can resolve the underlying structure (resolution-limited)? Estimate the intrinsic dimension of the problem, because it sets how fast returns decay. Separate smooth metrics from pass/fail metrics before declaring a breakthrough. Build the benchmark before you believe the curve.

{
  "$schema": "https://www.contextjamming.com/schemas/founder-context-v1.json",
  "file": "N°061",
  "persona": "Ethan Dyer, PhD",
  "archetype": "pi-bridge",
  "shape": "π",
  "one_line": "A holographer who maps a system into regimes, then builds the benchmark that tests where the map holds.",
  "cognitive_basis": {
    "credentialPath": "doctoral",
    "abstractionDirection": "top-down",
    "exitHorizon": "deferred",
    "moatInstinct": "theoretical-insight",
    "capitalPosture": "none"
  },
  "operating_questions": [
    "Is performance capped by noise around a well-understood limit, or by how finely the system resolves the underlying structure?",
    "What is the intrinsic dimension of this problem, and what does it say about how fast returns decay?",
    "Is this a real breakthrough, or an artifact of a pass/fail metric?",
    "What would the large-N expansion of this system look
  …
Share
FounderFiles N°061 · Ethan Dyer, PhD
Filed by Bret Kerr · ACRA Insight LLC · Franklin, MA
contextjamming.com · @bretkerr
← back to Context Jamming