ARTIFICIAL INTELLIGENCE · INTERACTIVE DOSSIER
When Intelligence Gets Cheaper: The Real Economics Behind Tokens
An interview proposes measuring AI like barrels of oil. Here we test the metaphor: comparable prices, non-comparable quality, cost per outcome, and privacy beyond a label.
On July 14, 2026, Chamath Palihapitiya brought a memorable image to CNBC: think of 1 million tokens as a “barrel of intelligence.” If two providers deliver something similar and one charges a fraction, competition should compress prices. The intuition is powerful because it turns an abstract bill into an everyday question: are we buying useful capability, or paying a premium our task does not need? It is also dangerous when read literally. A token does not contain a fixed quantity of intelligence, tokenizers count differently, and an input rate cannot simply be mixed with an output rate.
This dossier preserves the interview's thesis without inheriting its conclusions. We compare five official price lists current on 2026-07-17, reconstruct a common workload of 1M input and 0.2M output, and make every assumption visible. The result shows a wide cost spread, although not the same figures stated on air. That difference does not invalidate the economic question; it demonstrates why a serious comparison needs the exact model, token ratio, caching, context, tools, and cutoff date.
The second half changes the unit. Instead of asking what one million tokens cost, it asks what an accepted outcome costs. Repeated calls, retries, human review, latency, tools, errors, and privacy now appear. The lab does not predict your bill, and the chart does not rank quality. They are instruments for disciplined reasoning: define the task, measure it on your own distribution, and buy the minimum capability that keeps risk within explicit limits.
01 · BEGIN WITH THE THESIS
A useful metaphor that should not become a unit
The interview begins with an enterprise concern: several model companies are growing quickly, and someone in the chain must eventually turn that spend into profit. Palihapitiya compresses the tension into an oil comparison. If 1 million tokens were a barrel of intelligence, buyers should seek the cheapest barrel that moves their “engine” well enough. The metaphor makes competitive pressure visible and forces us to ask who captures value as capability diffuses.
The problem starts when the barrel is treated as a homogeneous commodity. A real barrel has a physical unit and specified grades; one million tokens merely counts processed fragments. The same document can yield different counts under different tokenizers. Input, output, reasoning, audio, images, and tools have separate rates. Even within one provider, long context, priority, or batch can change the price. The unit's name stays the same while the economic product changes.
There is a deeper difference: tokens are an input, not an outcome. One model may need fewer calls to solve a hard task, produce an answer requiring less review, or avoid an expensive error. Another may be cheap per call and fail exactly where the workflow needs precision. Without a definition of success, the buyer sees only the bill's numerator and loses the denominator that gives the purchase meaning.
This article therefore turns the assertion into a sequence of verifiable questions. Which exact model? What input/output ratio? Which date and service mode? How many iterations per task? What quality is accepted? What happens to the data? The thesis stops being “the price must fall” and becomes more useful: when different models compete for the same task, value shifts toward whoever proves the lowest total cost within the permitted risk.
02 · BUILD THE UNITS
What an API rate actually buys
Before comparing providers, draw the billing chain. A request carries instructions, context, and data as input. The model generates output and, depending on the product, also consumes internal reasoning. When it calls a tool, the result returns as new context and may trigger another loop. An agent interface adds memory, search, execution, storage, and observability. The published price per million tokens covers only specific parts of that chain.
The input/output split matters because their rates are often asymmetric. A classification application may read a lot and answer briefly; a report generator may produce long output; an agent may alternate between both for several rounds. Changing the ratio changes the bill and, in some cases, the relative order between providers. That is why this dossier's initial comparison keeps a visible 5:1 ratio and lets you explore others without calling them market averages.
The final unit should belong to the work, not the infrastructure. In support, it might be a case resolved without reopening; in extraction, a correct record; in programming, a change that passes tests and review; in research, a traceable claim that survives counterevidence. Only then can model cost be divided by accepted outcomes and compared with review, latency, or expected harm.
Billable token
Measures volume processed under provider rules; it contains no fixed capability measure.
Call
Groups one interaction's input and output, but a job may require several.
Task
Describes the work the system attempts to complete for a person or process.
Accepted outcome
Passes defined usefulness, quality, safety, and evidence criteria.
Escaped error
A failure that passes controls and creates rework, harm, or exposure.
Value
Benefit attributable to the outcome after all relevant costs and risks.
03 · NORMALIZE THE PRICE
Five official lists under the same workload
To test the metaphor, we use models with public rates and exact names: GPT-5.6 Sol, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, and DeepSeek V4 Pro. The initial view bills 1M input and 0.2M output without caching or tools. With rates observed on 2026-07-17, the derived totals are $11.00, $4.00, $3.30, $3.20, and $0.609. The spread is real and large enough to justify economic evaluation.
The comparison also shows why one figure can mislead. Claude Sonnet 5 has a promotional rate with an end date. Grok 4.5 changes rates after crossing the long-context threshold. DeepSeek V4 Pro distinguishes cache hit from cache miss. Google includes reasoning in output and bills grounding tools separately. OpenAI separates processing modes, caching, and tools. A “barrel” can omit the very component that dominates your bill.
Use the buttons to change the output ratio. Watch which part of each bar is input and which is output. This movement does not simulate users, estimate demand, or predict quality. It answers one narrow question: if we sent this exact volume under these list rules, what would the base charge be? Narrowness is a virtue because it prevents the chart from claiming more than the data supports.
The responsible next step is not to choose the shortest bar. It is to take two or three candidates into an evaluation of the real workflow, measure accepted outcomes, and register the features that change the contract. Public price forms a hypothesis; first-party telemetry decides whether that hypothesis survives.
The price frontier: same workload, different bills
The same token volume across five price lists. This normalizes volume; it does not make the models equivalent products.
It excludes caching, long context, tools, discounts, tax, latency, reliability, privacy, and success rate.
input × input rate + output × output ratePrice cutoff: 2026-07-17Open equivalent table, assumptions, and limitations
- The initial view uses 1M input tokens and 0.2M output tokens with Standard or equivalent rates published on 2026-07-17.
- DeepSeek V4 Pro uses cache-miss input, and Claude Sonnet 5 uses the promotional rate through August 31, 2026.
- The other workloads only change the input/output ratio; they are not an observed production distribution.
- It does not compare quality, latency, availability, safety, privacy, tokenization, or success rate.
- It excludes caching, long context, batch, priority, tools, grounding, tax, discounts, and private contracts.
- Prices and models change; the cutoff date is inseparable from the visualization.
| Provider | Model | Input / 1M | Output / 1M | total · 1M input · 0.2M output |
|---|---|---|---|---|
| OpenAI | GPT-5.6 Sol | $5.00 | $30.00 | $11.00 |
| Anthropic | Claude Sonnet 5 | $2.00 | $10.00 | $4.00 |
| Gemini 3.5 Flash | $1.50 | $9.00 | $3.30 | |
| xAI | Grok 4.5 | $2.00 | $6.00 | $3.20 |
| DeepSeek | DeepSeek V4 Pro | $0.435 | $0.87 | $0.609 |
04 · MEASURE CAPABILITY
Convergence does not mean equivalence
The interview argues that new models arrive at a fraction of the price while retaining much of the quality. Aggregate convergence has evidence: in March 2026, Anthropic 1503, xAI 1495, Google 1494, and OpenAI 1481 were separated by 22 Arena Elo points. That changes bargaining power and makes model substitution more plausible for common tasks. It does not prove that every model is 80% or 95% as good on every job.
A leaderboard compresses many interactions into one number. Compression is useful for general orientation and dangerous for local decisions. Arena reflects human preferences inside a platform; a coding benchmark reflects specific repositories and criteria; a reasoning test may contain invalid questions; a static evaluation can be trained or adapted against. NIST distinguishes benchmark accuracy from generalized accuracy because moving from one to the other requires assumptions that often remain hidden.
AI is also jagged. A system can solve advanced mathematics and fail a trivial instruction, produce excellent code and expose a secret, or summarize well in English while losing nuance in Spanish. The mean hides risk tails. A purchasing evaluation therefore needs normal cases, edge cases, attacks, sensitive data, long inputs, and tool failures. It also needs repetition: one brilliant demonstration does not estimate an error rate.
The practical question is “good enough for what, under which control, and at what failure cost?” An economical model may dominate low-risk classification and be rejected for security review; a premium model may justify its rate if it reduces human escalation or avoids a material error. Relevant quality does not live in a brand name: it lives in a task distribution, a threshold, and a reproducible measurement.
Representativeness
Cases should resemble the work, languages, data, and errors that will reach production.
Predefined criterion
Define what counts as correct before seeing which model wins.
Repetition
Measure variability instead of trusting one favorable output.
Risk tails
Include rare cases whose harm exceeds the average saving.
Complete system
Evaluate tools, retrieval, permissions, and review—not isolated text alone.
Drift
Version the model, prompt, data, and date to detect later change.
05 · SEE THE SYSTEM
Model price occupies only one layer
Token obsession appears when an organization measures what a provider bills instead of what a user receives. The interactive stack corrects the viewpoint. It begins with demand: someone needs an outcome at some frequency. Orchestration determines context, calls, and retries. The model contributes capability. Operations connect data and tools. Assurance detects failures. Only at the end does an accepted outcome appear.
Each layer can dominate cost for different reasons. In a massive classifier, the API bill may matter greatly. In a regulated process, human review may exceed it. In an agent, tool calls, repeated context, and failed retrievals can multiply tokens. In a low-latency application, reserved capacity or priority pricing may matter more than the Standard rate. Reducing one layer without observing the others displaces the problem.
The 3D view encodes no magnitude; it organizes dependencies. Select a layer and read its control question, minimum metric, and common failure. Then switch to 2D to verify that the information does not depend on the visual effect. The flow is also a circuit: outcome errors should return to evaluation, evaluation should change orchestration or model selection, and the new version should be measured again.
This structure explains why an API price war can coexist with rising total spend. If cheaper capability unlocks more tasks or longer agents, volume can expand. The company that records only tokens will see adoption; the company that records cost per outcome can distinguish useful expansion from consumption without convergence.
The value stack: from token to outcome
Rotate the system economics conceptually: model price is one layer, not the complete product.
Open equivalent table, assumptions, and limitations
- The six layers are an editorial reasoning taxonomy, not a mandatory architecture.
- The order represents conceptual dependencies: demand activates the flow, and the outcome feeds evaluation back.
- Thickness, color, and 3D separation encode no quantitative magnitude.
- It does not represent every technical, contractual, or regulatory component in a production system.
- One layer can run across several providers, and one isolated metric can create unwanted incentives.
- The 3D view is illustrative; the 2D view and table preserve the same information.
| Layer | Question | Minimum metric |
|---|---|---|
| Demand | What outcome does someone need? | Useful tasks / month |
| Orchestration | How many loops are needed? | Calls and tokens / task |
| Model | What minimum capability works? | API USD / task |
| Operations | Which services sustain the flow? | USD and p95 / workflow |
| Assurance | How do we catch errors? | Minutes and escaped errors |
| Outcome | Was the job solved? | USD / accepted outcome |
06 · CALCULATE THE FLOW
From unit price to cost per accepted outcome
The lab begins with 10,000 monthly tasks. Each task makes 2 calls, each call uses 5,000 input tokens and 1,000 output tokens, and a 10% retry rate raises the billable count. The organization accepts 90% of outcomes and reviews each task for 3 minutes at $30 per hour. These values are a pedagogical scenario, not an enterprise statistic. Their purpose is to expose the formula and let you break it.
First move tasks or calls. API cost grows linearly because it multiplies volume. Next increase retries and watch a small rate affect every call. Change the model to isolate the rate difference, but do not read the comparison as substitution: the lab artificially holds acceptance constant. In a real decision, every candidate needs its measured rate and perhaps different review minutes.
Human review reveals something token price hides. If a more capable model costs more in API charges but reduces errors, escalation, or inspection time, it can produce a cheaper outcome. The opposite also happens: a mandatory review layer may dominate so completely that optimizing cents of inference barely moves the total. The interface shows the API share so the discussion can focus on the material variable.
The key result is operating cost divided by accepted outcomes. Even so, it is not ROI. Tools, storage, network, support, implementation, tax, discounts, revenue, and escaped-error cost are missing. Use the reflection field to write the metrics you would measure before migrating. If you cannot name them, you do not yet have a value comparison—only a rate comparison.
Cost-per-accepted-outcome lab
Change one variable at a time. The lab separates the model bill from the cost of review and accepted outcomes.
Same workload, API bill only
Keeps tokens and calls fixed; it does not assume equal quality or acceptance rates across models.
- OpenAIGPT-5.6 Sol$1,210.00
- AnthropicClaude Sonnet 5$440.00
- GoogleGemini 3.5 Flash$363.00
- xAIGrok 4.5$352.00
- DeepSeekDeepSeek V4 Pro$66.99
spend = calls × tokens × rate + human reviewOpen formulas, equivalent table, assumptions, and limitations
- The user defines tasks, calls, tokens, retries, acceptance, and review; the initial scenario is pedagogical, not an enterprise average.
- Human review is modeled as minutes per task multiplied by a loaded hourly cost.
- The provider comparison holds the exact workload and acceptance rate fixed only to isolate the API bill.
- It excludes tools, storage, retrieval, network, observability, support, tax, and discounts.
- Acceptance rate must be measured with first-party evaluations; holding it fixed across models does not imply equivalence.
- It does not predict ROI, revenue, productivity, or errors for a real organization.
| Variable | Initial value | Role in the formula |
|---|---|---|
| Tasks | 10,000 / month | Multiplies usage and review |
| Calls | 2 / task | Multiplies tokens |
| Input | 5,000 / call | Billed at input rate |
| Output | 1,000 / call | Billed at output rate |
| Retries | 10% | Increases billable calls |
| Acceptance | 90% | Defines accepted outcomes |
| Review | 3 min at $30/h | Adds human cost |
| Cost per accepted outcome | $1.80 / outcome | $16,210 ÷ 9,000 accepted outcomes |
| Metric | Value | Unit and basis |
|---|---|---|
| Selected model | OpenAI · GPT-5.6 Sol | SRC-303 |
| Input rate | $5.00 | USD / 1M input tokens |
| Output rate | $30.00 | USD / 1M output tokens |
| Tasks per month | 10,000 | tasks / month |
| Calls per task | 2 | calls / task |
| Input tokens per call | 5,000 | tokens / call |
| Output tokens per call | 1,000 | tokens / call |
| Additional retries | 10% | % extra over base calls |
| Accepted outcomes | 90% | % of tasks accepted |
| Human review per task | 3 | min / task |
| Reviewer hourly cost | $30.00 | USD / hour |
| billable calls | 22,000 | billable calls / month |
| Model bill | $1,210.00 | USD / month |
| Human review | $15,000.00 | USD / month |
| Scenario operating cost | $16,210.00 | USD / month |
| Accepted outcomes | 9,000 | accepted outcomes / month |
| Cost per accepted outcome | $1.80 | USD / accepted outcome |
07 · FOLLOW THE DATA
ZDR does not fit in a yes-or-no box
The interview is right to distrust a superficial reading of Zero Data Retention, but its “like” button example mixes possibilities without separating products. Current documentation supports a more precise explanation. OpenAI, Anthropic, and Google describe controls for API customers or paid services, but coverage depends on eligibility, contract, organization, endpoint, model, feature, and configuration. It should not be automatically extrapolated to a consumer interface.
Training and retention must be separated. “We do not train on your content by default” limits one purpose. ZDR aims to prevent prompts and responses from being stored at rest after the response under covered conditions. A provider can offer the first restriction without every feature satisfying the second. Files, batches, persistent sessions, execution, caching, feedback, or abuse monitoring can have different lifecycles.
Application state exists too. A responses system can save conversations if requested; a files API must retain a file to use it; a managed agent needs a session; an external tool receives its own data slice. Model ZDR does not delete copies in your database, observability stack, proxy, search provider, or ticketing system. The privacy boundary crosses the entire flow.
Metadata deserves a separate column. Account identifiers, billing, usage, latency, errors, safety filters, or network addresses may be processed to operate the service even when content has stricter controls. That does not make ZDR useless; it makes it one concrete control inside an architecture. Proper verification records which data, who processes it, where, for what purpose, for how long, and under which exception.
Before sending secrets, personal information, or regulated data, a team should review the applicable contract and configuration, minimize content, isolate credentials, limit tools, define deletion, and test logs. This dossier is educational and does not replace legal, security, or compliance review. Its safest recommendation is methodological: do not trust the control's name; verify each data path.
Training
Can content be used to improve models, and what is the default?
Content retention
Are prompts and responses stored, for how long, and under which exceptions?
Application state
Do files, conversations, batches, caches, or sessions persist to function?
Metadata
Which usage, safety, billing, and network data is processed apart from content?
Third parties
Which tools, clouds, or integrations receive data, and under what terms?
Your infrastructure
Which copies remain in proxies, logs, databases, queues, evaluations, and support systems?
08 · BUY WITH EVIDENCE
A process that survives the next price drop
The model names and rates in this article will become stale. A good process will not. Start by describing one task with inputs, outputs, volume, users, and possible harm. Build a case set that includes language, long data, tool errors, and rare situations. Define criteria before execution: accuracy, format, evidence, safety, latency, and when the system must escalate to a person.
Then run several candidates with versioned configuration. Record tokens by modality, calls, tools, time, errors, acceptance, and review. Calculate cost per accepted outcome and segment by difficulty; an average can hide that the economical model fails in the most valuable segment. For heterogeneous tasks, test routing: a cheap model handles routine work while a premium model receives difficult or high-risk cases.
Add operating limits. Per-task budgets, output ceilings, iteration caps, timeouts, controlled caching, and alerts prevent a capability improvement from becoming open-ended consumption. Repeat the evaluation whenever the model, prompt, tools, data, or policy changes. The exact date and identifier are part of the result, not administrative details.
Finally, treat privacy and continuity as purchasing criteria. Verify contract, region, retention, export, deletion, dependency on proprietary features, and exit plan. The cheapest token can be expensive if it blocks migration, lacks reserved capacity, or forces controls to be rebuilt. The premium option can be unjustifiable if it shows no improvement in the denominator.
1 · Define
Task, user, volume, sensitive data, and error cost.
2 · Sample
Normal, difficult, rare, adversarial, and multilingual cases.
3 · Accept
Predefined quality, safety, latency, and escalation criteria.
4 · Measure
Outcomes, tokens, calls, tools, review, and variability.
5 · Calculate
Total cost per accepted outcome and per risk segment.
6 · Control
Budgets, limits, observability, permissions, and alerts.
7 · Revalidate
A new test after any model, data, or workflow change.
09 · RETURN TO THE INTERVIEW
What remains after testing the thesis
The core intuition about price pressure survives. Wide list-price differences exist, and aggregate convergence lets more buyers test alternatives. A model brand's value cannot rest solely on arriving first. It must be sustained by measurable quality, reliability, features, distribution, security, support, or superior total economics. The question “why pay more?” is healthy when it forces a provider to demonstrate the denominator.
The specific “barrel” figure does not survive as a universal comparison. Current official prices depend on model, input, output, tokenizer, and conditions. A mixture without a visible formula cannot be audited. Nor is the claim that economical models retain 80% or 95% of quality for most uses demonstrated. Arena provides a convergence signal, not a universal substitution rate.
The warning about enterprise spend is possible, not observed in the interview. An organization may discover ungoverned consumption, but it may also negotiate capacity, use subscriptions, receive discounts, or generate benefits above cost. The correct scenario is not to predict corporate earnings surprises; it is to install telemetry before the bill grows: budgets, cost per task, retries, review, and accepted value.
The hardware and memory thesis belongs to investing. The AI Index documents rising compute spend, and infrastructure may remain scarce, but that does not determine margins, multiples, or returns for a specific company. This article recommends no securities and does not validate that one segment will outperform the market. Separating physical demand from financial return keeps a technical explanation from becoming implicit advice.
The privacy observation deserves attention with a better taxonomy. ZDR is real and useful within its coverage; it is not a magic layer controlling consumers, tools, and external copies. The interview's most robust conclusion is not that all AI becomes a commodity. It is that, as capability gets cheaper, advantage shifts toward designing, measuring, and governing the system that turns capability into trustworthy outcomes.
CHECK TRANSFER
Can you decide without the metaphor?
Answer through reasoning. Every explanation identifies the distinction that matters in production.
Traceability
Evidence ledger
CLM-301Chamath Palihapitiya proposes calling 1 million tokens a “barrel of intelligence” and argues that the price gap will have to rationalize.direct
CLM-302GPT-5.6 Sol, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, and DeepSeek V4 Pro publish separate input and output rates, with additional rules for caching, context, tools, or batch processing.triangulated
CLM-303With 1 million input tokens and 0.2 million output tokens, without caching or tools, the calculated list cost is $11.00, $4.00, $3.30, $3.20, and $0.609, respectively.derived
Locator: Reproducible derivation: input price × 1 + output price × 0.2; DeepSeek V4 Pro cache-miss rate and current Claude Sonnet 5 promotional rate
Uncertainty: The comparison normalizes volume, not quality, latency, reliability, safety, tokenization, availability, or total task cost.
Sources and limitations: SRC-303SRC-304SRC-305SRC-306SRC-307
CLM-304A token price does not measure task quality: evaluation requires declaring the measurement target, cases, assumptions, and uncertainty.triangulated
CLM-305In March 2026, Anthropic 1503, xAI 1495, Google 1494, and OpenAI 1481 were separated by 22 points in Arena Elo.direct
Locator: AI Index 2026 Technical Performance, finding 2
Uncertainty: Arena aggregates human preferences, and its ranking may reflect adaptation to the platform; it is not an evaluation of every enterprise workflow.
Sources and limitations: SRC-308
CLM-306Aggregate convergence increases competitive pressure, but it does not prove that a cheap model is interchangeable with a premium model on a specific task.derived
CLM-307The billable cost of an agent can include input, output, and reasoning tokens, cache reads or writes, tool calls, and multiple iterations.triangulated
CLM-308Reducing price per token does not guarantee lower total spend if volume, repeated context, iterations, or retry rates grow faster.derived
Locator: Simulator's reproducible arithmetic identity: spend = billable units × rate; direction depends on both factors
Uncertainty: This describes a mathematical condition; it does not claim every company is increasing usage or predict their results.
Sources and limitations: SRC-303SRC-304SRC-305SRC-306SRC-307
CLM-309Zero Data Retention is a contractual, organizational, endpoint, and feature property; by itself it does not mean the absence of metadata, application state, or third parties.triangulated
Locator: OpenAI ZDR eligibility table; Anthropic What ZDR does not cover and Feature eligibility; Gemini achieving zero data retention
Uncertainty: Specific contracts and configurations govern over general documentation and may differ by region, product, or account.
CLM-310Not using content for training and not retaining content after the response are different controls; a service can offer one without automatically guaranteeing the other.triangulated
CLM-311Rational AI procurement optimizes cost per accepted outcome, not tokens consumed or an isolated leaderboard position.derived
CLM-312The interview argues that hardware and memory will keep benefiting from scarcity, but that is an investment thesis, not a conclusion proven by API prices.disputed
CLM-313The CNBC Squawk Box interview aired on July 14, 2026 and introduces Palihapitiya as founder and CEO of Social Capital, CEO of 8090, and an All-In host.direct
CLM-314The visual comparison uses a reference workload of 1M input and 0.2M output, an illustrative 5:1 ratio that does not describe every application.derived
Operational bibliography
Sources and limitations
- SRC-301internal artifactInternal evidence · not publicly available
All-In podcast host Chamath Palihapitiya on the current state of AI — supplied transcript
Transcript of the CNBC interview supplied by the user · 2026-07-17
Locator: Full transcript with 00:01–13:18 timestamps; preserved bytes and manifests under content/intake/economics-of-cheap-intelligence
Integrity fingerprint:
This is an automated transcript that ends mid-sentence; it preserves the interviewee's opinions but does not validate his figures or forecasts.sha256-682881672216509ee1142c4e40163c3ccc3f6c8b9901ed944e09c0aa5f964c48 - SRC-302primary source
All-In podcast host Chamath Palihapitiya on the current state of AI
CNBC Squawk Box · 2026-07-17
Locator: July 14, 2026 video; segments 01:42–03:59, 05:31–08:12, and 12:13–13:18
The interview presents the guest's thesis and analogies; its prices mix models, input/output ratios, and possibly unstated assumptions. - SRC-303primary source
Models — OpenAI API
OpenAI · 2026-07-17
Locator: Frontier models section: GPT-5.6 Sol, Terra, and Luna; input and output prices per MTok
List prices may change; they exclude tools, priority, batch, caching, tax, discounts, and private contracts. - SRC-304primary source
Precios — Claude Platform Docs
Anthropic · 2026-07-17
Locator: Model pricing table, Claude Sonnet 5 and Claude Sonnet 4.6 rows; caching, batch, and tool sections
The Claude Sonnet 5 rate used in the article is promotional through August 31, 2026; tokenization also differs across generations. - SRC-305primary source
Gemini Developer API pricing
Google AI for Developers · 2026-07-17
Locator: Gemini 3.5 Flash and Gemini 3.1 Flash-Lite sections; Standard, Batch, and Context caching tables
Output includes thinking tokens; tools and grounding are billed separately, and Preview models may change. - SRC-306primary source
Pricing — xAI Docs
xAI Corp. · 2026-07-17
Locator: Text API, Prices per 1M tokens table, grok-4.5 row; tools, batch, and priority sections
The rate changes for long context; server tools add charges, and availability depends on the account. - SRC-307primary source
Models & Pricing — DeepSeek API Docs
DeepSeek · 2026-07-17
Locator: Model Details: deepseek-v4-flash and deepseek-v4-pro; cache-hit, cache-miss, and output prices
The article uses the V4 Pro cache-miss rate; the provider warns prices may change, and its metrics do not make models equivalent. - SRC-308secondary source
Technical Performance — 2026 AI Index Report
Stanford Institute for Human-Centered AI · 2026-07-17
Locator: Findings 2, 5, 7, 8, and 9: Arena convergence, benchmark fragility, and jagged intelligence
This is a synthesis of heterogeneous evaluations; Arena Elo reflects aggregate preferences and does not guarantee performance, safety, or cost on a business task. - SRC-309primary source
Expanding the AI Evaluation Toolbox with Statistical Models — NIST AI 800-3
National Institute of Standards and Technology · 2026-07-17
Locator: Abstract and contributions: benchmark accuracy, generalized accuracy, assumptions, and uncertainty
The framework improves statistical interpretation but does not prescribe one universal benchmark, provider, or threshold for every business. - SRC-310primary source
Data controls in the OpenAI platform
OpenAI · 2026-07-17
Locator: Modified Abuse Monitoring and Zero Data Retention sections; eligibility and application-state table
Eligibility requires approval, and some capabilities retain application state even with ZDR; current contracts are the final source. - SRC-311primary source
API and data retention — Claude Platform Docs
Anthropic · 2026-07-17
Locator: How Anthropic approaches data retention, Zero data retention, What ZDR does not cover, and Feature eligibility sections
ZDR is enabled per organization and does not cover every feature, consumer product, interface, or integration; some models require retention. - SRC-312primary source
Zero data retention in the Gemini Developer API
Google AI for Developers · 2026-07-17
Locator: Training restriction and Customer data retention and achieving zero data retention sections
Paid services restrict training, but achieving ZDR requires configurations and avoiding retaining features; operational data also remains. - SRC-313secondary source
Economy — 2026 AI Index Report
Stanford Institute for Human-Centered AI · 2026-07-17
Locator: Section 4.2 Investment and Infrastructure, Figure 4.2.20, PDF page 190
Annual compute spend is an Epoch AI estimate and proxies rented capacity; it does not prove future profitability for labs or hardware makers.