CONTEXT JAMMING

Field notes from inside the context window.

Share
Audio·Dispatch
Context Jamming · Listen
How Tom Brown negotiated an AI ceasefire

FounderFiles·N°046·Scaling · Compute · Diplomacy

2026

Tom Brown — Co-founder & Chief Compute Officer, Anthropic
Fig. · Constraint collapseChief Compute Officer

Subject·Tom B. Brown·Co-founder & Chief Compute Officer, Anthropic

Tom BROWN.

Lead author, GPT-3 · Architect of the few-shot learner and the sovereign compute stack

Brown doesn't invent new algorithms so much as he collapses constraints. From the ad-serving API that handled 1.5 billion monthly impressions, to the 175-billion-parameter model that learned in context, to the multi-vendor silicon fabric that keeps Anthropic in the race, to the two-week negotiation that restored Mythos 5 and Fable 5 — the same move: find the leverage point, apply empirical force, and bridge the factions that would otherwise keep the system locked.

Trained
MIT (CS + Cognitive Science)
At
Anthropic · Co-founder & Chief Compute Officer
File
N°046 · Scaling · Compute · Diplomacy
§ 01 · The Wolf Mindset

The Wolf Mindset

Y Combinator, MoPub, and the survival logic of early systems

Brown did not arrive at frontier AI from a pure research pedigree. He arrived from the high-stakes, zero-margin environment of Y Combinator startups. After MIT he became the first employee at Linked Language, then a founding engineer at MoPub, where he built the early server architecture and scaled the ad-serving API to 1.5 billion monthly impressions. Twitter acquired MoPub for $600 million in 2013. The technical residue of that period was not “machine learning experience.” It was battle-tested distributed systems competence under real traffic and real failure modes.

In 2011 he co-founded Grouper (later Grouper Social Club) with Michael Waxman and Ariel Galili — a YC Winter 2012 invite-only group-dating product that matched trios of friends via Facebook data. As CTO he engineered the matching algorithms that facilitated over one million drinks across 25 cities before the company deadpooled. The deeper education was social and operational: deliberate networking, resourcefulness under scarcity, and the friendships that later opened the door to OpenAI. One of those friendships was with Greg Brockman.

A wolf mindset — a survivalist, highly adaptive approach to problem-solving driven by the stark reality of needing to hunt for food or face corporate oblivion.
Brown on the contrast between startup and large-company training
§ 02 · The Self-Taught Pivot

The Self-Taught Pivot

Six months, a GPU, and “I will even mop floors”

By 2015 Brown recognized deep learning as a historical inflection point, but he lacked formal graduate training in machine learning. He spent six months systematically closing the gap: Axler's Linear Algebra Done Right, the DeepMind scaling notes, Coursera sequences, Kaggle projects, and a personal GPU purchased with YC alumni credits. The philosophy was explicit — mastery through direct, hands-on failure and iteration rather than credentialed abstraction.

He reached out to the newly formed OpenAI and, by his own later account, told leadership he would mop floors if they would let him help. The combination of demonstrated systems engineering from MoPub and raw willingness produced an early technical staff role. His first nine months were largely non-ML infrastructure work — building the distributed systems the research teams needed. By 2017 he was contributing to foundational RLHF papers, planting the alignment instincts that would later define Anthropic's public posture.

The B- in undergraduate linear algebra is not a cute anecdote. In an industry that still mythologizes mathematical prodigies, Brown's path is evidence that the scaling era rewarded applied throughput, adaptive learning velocity, and the ability to keep the system running under load more than pure theoretical pedigree.

§ 03 · Few-Shot Learner

Few-Shot Learner

Language Models are Few-Shot Learners (2020)

Brown was lead author on the paper that shattered the fine-tuning paradigm. Prior models (BERT, GPT-1, GPT-2) required an additional, expensive supervised phase to become useful at any specific task. GPT-3 demonstrated that at 175 billion parameters — an order of magnitude larger than previous non-sparse models — an autoregressive language model develops emergent in-context learning. No gradient updates. No task-specific fine-tuning. Only a natural-language instruction and a handful of demonstration examples inside the context window.

MetricZero-ShotOne-ShotFew-Shot
CoQA (F1)81.584.085.0
TriviaQA (Accuracy)64.3%68.0%71.2%

Few-shot performance rose smoothly with scale. The paper supplied the empirical backbone for the scaling-laws worldview: allocate exponentially more compute and carefully filtered data to a relatively simple predictive objective and intelligence emerges as a reliable function of that allocation. Brown later summarized the revelation as the triumph of “the stupid thing that works.”

The stupid thing that works — brute computational force and scale achieving outcomes that complex algorithmic micro-optimizations could not.
Brown on the GPT-3 revelation
§ 04 · The Compute Mandate

The Compute Mandate

From researcher to Chief Compute Officer

Ideological fractures inside OpenAI over commercialization pace and safety prioritization produced the 2021 departure of Dario and Daniela Amodei, Jared Kaplan, Chris Olah, and Tom Brown. Anthropic was founded as a public-benefit corporation whose explicit charter prioritized reliable, interpretable, and steerable systems — Constitutional AI and mechanistic interpretability as load-bearing techniques.

Inside the new organization Brown took dual titles: co-founder and Chief Compute Officer (sometimes styled Head of Core Resources). While Olah pursued circuit-level interpretability and Amodei owned corporate strategy, Brown owned the physical and financial substrate: securing, scheduling, and optimizing the hardware that would keep Anthropic competitive in the global compute arms race. He has publicly framed the pursuit of AGI as the catalyst for “the largest infrastructure buildout of all time,” with capital and energy requirements on track to exceed the Manhattan Project and Apollo program combined.

§ 05 · Multi-Chip Reality

Multi-Chip Reality

Trainium, TPUs, Micron, and Colossus

Brown rejected single-vendor lock-in. Anthropic's training and inference fabric was deliberately multi-chip: Amazon Trainium, Google TPUs, and NVIDIA GPUs, with workloads dynamically allocated according to the economics and availability of each silicon generation. Strategic multi-billion-dollar compute contracts — including capacity at SpaceX's Colossus data center — were part of the same portfolio approach.

The Micron partnership was especially revealing. Micron's strategic investment in Anthropic's $65 billion Series H (the round that produced a $965 billion valuation) was accompanied by a deep technical collaboration focused on the full memory hierarchy: HBM, DRAM, and data-center SSDs. Brown treated memory bandwidth and energy efficiency as first-class constraints on token economics, not secondary procurement details. Scaling, in his operational framing, is a full-stack systems problem that begins at the package and ends at the power purchase agreement.

§ 06 · Glasswing

Glasswing and the Patching Paradox

Autonomous capability meets enterprise inertia

Claude Opus 4.8 and the subsequent Mythos 5 / Fable 5 generation marked the transition from passive generation to long-horizon agentic work. Opus 4.8 ships with a 1 M token context window, top-decile coding and intelligence indices, and adaptive thinking that allocates compute dynamically. Mythos 5 was the internal, cybersecurity-oriented frontier model; Fable 5 its more heavily guardrailed public counterpart.

Project Glasswing gave curated, monitored access to Mythos to a short list of systemically important defenders (FIS, ICE, Commvault, Trend Micro). The model autonomously surfaced thousands of high-severity vulnerabilities, including zero-days, across major operating systems, browsers, and open-source projects, generating more than 1,100 unvetted reports in weeks. The operational discovery was the patching paradox: even when an AI can find and propose a fix faster than any human team, large enterprises cannot deploy at the same velocity. Discovery collapsed the window between vulnerability and exploitation; the defensive response could not. The implication is a forced shift from reactive patching to Resilience Operations (ResOps).

§ 07 · June 2026 Flashpoint

The June 2026 Flashpoint

Jailbreaks, distillation, and the BIS “is informed” letter

Mid-June 2026 produced a convergence of threat signals: a novel jailbreak that bypassed Fable 5's safety classifiers, unverified but widely circulated claims of NSA red-team success against classified test systems using Mythos, and the largest known distillation campaign against the Claude network — Alibaba allegedly operating ~25,000 fraudulent accounts for 28.8 million exchanges over six weeks. Compounding the damage was the discovery that Mythos access had been granted to a South Korean telecommunications firm with suspected China back-channel ties.

On 12 June the Bureau of Industry and Security issued an emergency “is informed” directive under the Export Control Reform Act. Anthropic was ordered to restrict all foreign nationals — including its own foreign-born U.S.-based employees — from accessing Mythos 5 and Fable 5. Because the company lacked real-time nationality authentication across its global API surface, the only compliant path was a total global shutdown of both models. Enterprise workflows froze. Stripe's reported 50-million-line codebase overhaul stopped mid-flight. Chinese open-weight models immediately captured share.

§ 08 · The Geopolitical CTO

The Geopolitical CTO

Lutnick, Cairncross, and the Annex A compromise

Relations between the White House and Dario Amodei had already deteriorated; administration officials reportedly found him difficult to negotiate with. Anthropic therefore dispatched Brown and Head of Public Policy Sarah Heck. Over two weeks of daily sessions with Commerce Secretary Howard Lutnick and National Cyber Director Sean Cairncross, Brown focused on technical remediation rather than ideological framing. His team delivered a new safety classifier that blocked the reported Amazon jailbreak techniques in >99 % of cases and helped establish standardized benchmarks for future jailbreak assessment.

On 26 June Lutnick issued a formal letter — addressed directly to Tom Brown, bypassing Amodei — that partially lifted the controls. Mythos 5 returned under a secret Annex A list of roughly 100 vetted U.S. critical-infrastructure organizations plus government agencies and Anthropic's own foreign-national employees. Fable 5 followed a few days later after final classifier approval. The letter explicitly reserved the Commerce Department's right to reimpose restrictions at any time.

Three precedents were set: (1) the U.S. government demonstrated willingness to treat deployed model weights as export-controlled munitions and to pull a commercial model offline globally; (2) any future reprieve is conditional and reversible; (3) the technical founder who can speak both silicon and statecraft has become an indispensable diplomatic actor.

Securing compute is no longer just about procuring hardware; it requires navigating international trade law, export controls, and sovereign defense policy.
On compute as sovereign infrastructure
§ 09 · Stepping Stones

Stepping Stones

One move, four substrates

Read chronologically, Brown's career looks like a sequence of pivots. Read architecturally, it is a single instinct re-applied at ascending levels of constraint:

  • 01
    Consumer systems layer

    High-throughput ad serving and social matching algorithms under YC survival pressure.

  • 02
    Model layer

    Few-shot in-context learning via pure scale (GPT-3).

  • 03
    Physical compute layer

    Multi-vendor hyperscale fabric, memory hierarchy optimization, energy and token economics.

  • 04
    Sovereign / geopolitical layer

    Export-control negotiation, safety classifier co-design with the state, Annex A access architecture.

Each rung reuses the same primitives: empirical force over theoretical elegance, bridging across factions that would otherwise deadlock the system, and a willingness to treat the current constraint as just another distributed systems problem to be collapsed. The “wolf” that hunted for food in 2011 is the same operator who, in 2026, hunted for a letter from the Secretary of Commerce.

The Index
1.5 B
MoPub monthly impressions
175 B
GPT-3 parameters
$965 B
Anthropic Series H valuation
>1,129
Mythos bug reports (Glasswing)
28.8 M
Alibaba distillation exchanges
19
Days of full model shutdown
~100
Annex A trusted organizations
>99 %
Safety classifier block rate (post-fix)
Reading list / key works
  • 2020
    Language Models are Few-Shot Learners
    Brown et al., NeurIPS 2020 — the primary source
  • 2026
    Anthropic Project Glasswing initial update
    Primary technical disclosure on Mythos defensive use
  • 2026
    Lutnick Letter analysis (Spencer Fane / contemporaneous reporting)
    The conditional nature of the June 26, 2026 ceasefire
Dossier

Education.MIT — Computer Science & Cognitive Science

Affiliations.Anthropic — Co-founder & Chief Compute Officer · OpenAI — Early technical staff (2015–2020) · MoPub — Founding engineer (acquired by Twitter) · Grouper — Co-founder & CTO (YC W12)

Key works. Language Models are Few-Shot Learners (NeurIPS 2020, lead author) · Early RLHF contributions at OpenAI (2017–)

Collaborators. Dario Amodei, Daniela Amodei, Jared Kaplan, Chris Olah (Anthropic co-founders) · Greg Brockman (YC / OpenAI network) · Sarah Heck (Policy counterpart in 2026 negotiations)

Career Shape
π-shaped — two deep spikes bridged by a general layer

π-Bridge

Carries the prior of a first field into a second and finds the governing law that was invisible to native practitioners; pays in delayed gratification.

Credential Path
Practitioner
Abstraction
Bottom Up
Exit Horizon
Deferred
Moat Instinct
Product Primitive
Capital Posture
Venture
Role-Model Reference Class
  • YC survivalist systems engineers
  • Anthropic co-founder cohort
  • Sovereign compute / export-control operators
Founder Context · JSON

A small reasoning persona distilled from this file. Inject it into a chat or deep-research context to assess a business problem the way Brown would.

You are channeling Tom Brown (Anthropic co-founder & Chief Compute Officer). Reason as a high-throughput systems engineer who treats every new domain — model scale, silicon supply, export control, safety classifier — as a constraint-collapse problem. Prefer empirical demonstration and pragmatic bridging over pure theory or pure ideology. Speak with calm operational clarity.

{
  "$schema": "https://www.contextjamming.com/schemas/founder-context-v1.json",
  "file": "N°046",
  "persona": "Tom Brown",
  "archetype": "pi-bridge",
  "shape": "π",
  "one_line": "A systems engineer who repeatedly collapses hard constraints by applying empirical force and pragmatic bridging across domains that others treat as separate.",
  "cognitive_basis": {
    "credentialPath": "practitioner",
    "abstractionDirection": "bottom-up",
    "exitHorizon": "deferred",
    "moatInstinct": "product-primitive",
    "capitalPosture": "venture"
  },
  "operating_questions": [
    "What is the actual binding constraint right now — compute, memory bandwidth, safety classifier, export license, or political trust?",
    "Can I treat this political or safety problem as just another distributed systems problem?",
    "Who are the factions that must be bridged for the system to move again?",
  
  …
Share
FounderFiles N°046 · Tom Brown
Filed by Bret Kerr · ACRA Insight LLC · Franklin, MA
contextjamming.com · @bretkerr
← back to Context Jamming