Francesco Di Costanzo

(25) Local AI Will Make the Cloud More Valuable

Applied AI & Agent Systems
  • Local AI
  • Model economics
  • AI infrastructure
  • Enterprise AI

The $0.415 cloud call

On 25 August 2026, Perplexity launched Portable Computer, an agent system whose planner, tool router, scheduler, search index and 27-billion-parameter model run on an NVIDIA DGX Spark. It reads private files and executes ordinary workflows without consuming Perplexity credits. When the local model needs current information or stronger reasoning, it can ask the user to send a bounded subtask to one of more than fifteen cloud models.

Perplexity’s most revealing launch result came from Terminal-Bench 2.1, which tests agents on complex terminal tasks. The company reports that its local Qwen 3.8 27B configuration scored 59.6%. Adding Claude Opus 5 as an advisor raised the score to 73.0% at an estimated $0.415 per rollout. Running the frontier model alone scored 82.4% at about $0.65. The advisor received approved text, not direct access to local files or tools. These are vendor-run results from a system whose internal knowledge-work benchmark was not yet public at launch, so they are evidence of a mechanism rather than an independent verdict on the product.

The mechanism is economically important. Local inference can absorb private, repetitive and high-volume work, while the cloud handles the smaller set of steps where extra capability changes the outcome. My thesis is falsifiable: as local and open-weight use rises, frontier cloud models should process a smaller share of total tokens but retain a disproportionate share of spending and accepted high-stakes work. If local systems take spend as quickly as they take volume, the thesis fails.

Production traffic already has a barbell shape

Vercel’s AI Gateway provides an imperfect but useful view of production demand. It routes tens of trillions of tokens across hundreds of models, and its August 2026 index covers traffic through July. Open-weight models processed 36% of gateway tokens, up from 11% in April, while their share of spending reached 8.6%. DeepSeek alone ran a quarter of all tokens. Anthropic, by contrast, collected 65.1% of spending on 30% of volume, and its average token cost was 4.4 times the average for every other lab.

The split becomes sharper inside one workload. In coding agents, DeepSeek processed nearly a third of token volume while Anthropic collected more than four-fifths of spending. Back-office agents also consumed roughly two and a half times as much spending as their token share. Buyers were not replacing expensive intelligence with cheap intelligence across the board. They were assigning the two to different work.

Vercel’s sample is not the whole enterprise market. Gateway users are more likely than average buyers to operate multi-model applications, and Vercel values requests at published list prices rather than each customer’s negotiated bill. The direction is still hard to dismiss. In July, volume grew 59% and spending grew 37%, while the average price per token fell 13.6% because teams changed their model mix. Among teams that processed more than ten million tokens in both June and July, three quarters moved at least a tenth of their traffic to different models.

That is the production version of reserving premium capability for work that can use it. Cheaper systems expand the quantity of AI work that can be attempted. Frontier systems keep the work for which a failed answer, retry or human review costs more than the inference.

Local compute removes the frontier tax from routine work

A DGX Spark has 128 GB of unified memory and a regulatory maximum power draw of 233.2 watts. Its cost is paid before the first prompt; electricity, maintenance and operator time continue whether utilisation is high or low. Cloud APIs reverse that shape. They turn infrastructure into a variable charge, scale without a local capacity purchase and expose a catalogue that ranges from cheap classifiers to frontier reasoning.

Published prices show how wide that catalogue has become. OpenAI charges $0.20 per million input tokens and $1.20 per million output tokens for GPT-5.6 Luna, against $4 and $20 for GPT-5.6 Sol. Anthropic’s current range runs from $1 and $5 for Haiku 4.5 to $10 and $50 for Fable 5. Price-performance also keeps improving: OpenAI cut Luna’s price by 80% in July, while research using historical benchmark and price data estimates that the cost of a fixed level of model performance has fallen fivefold to tenfold per year across several task families.

Local ownership therefore wins on cost only under defined conditions. Demand must be steady enough to use the hardware, the local model must clear the task’s quality threshold, and the organisation must be able to operate the runtime. Research on owning private inference reaches the same boundary: privacy and control can justify the machine before token savings do. Sporadic demand, burst capacity and the hardest tasks still favour rental.

The surprising result is that removing routine cloud traffic can improve the cloud provider’s revenue mix. A frontier model no longer spends most of its capacity summarising files or reformatting tables. It receives the uncertain tail, where its incremental quality has a higher willingness to pay. Local AI reduces cloud volume while sharpening the frontier premium.

Routing becomes the economic control

Model routing has moved from a cost-saving trick to the control point of a hybrid system. RouteLLM, published at ICLR 2025, reduced evaluated cost by more than half without lowering response quality by learning when to choose a stronger model. FrugalGPT showed earlier that cascades could stop after a cheap model produced an acceptable answer. Later work added calibrated confidence, latency and capacity constraints.

The research also provides a warning. LLMRouterBench evaluated more than 400,000 instances across 33 models and found that several recent routing methods, including commercial systems, did not reliably beat a simple baseline. Larger model pools produced diminishing returns, and a substantial gap remained between practical routers and an oracle that always knew the right choice. A weak router either escalates too often and recreates the cloud bill, or escalates too late and accepts an avoidable failure.

Portable Computer uses several signals at the boundary. It keeps private documents and inference on the device, calls web search when freshness is required, and asks permission before content moves to a cloud service. Its reported advisor flow selects context, flags personal information and gives the remote model no tool authority. That design turns the cloud model into advice inside a locally controlled workflow.

The non-obvious product is the escalation policy. Cloud inference becomes an option on uncertainty: the buyer pays when the expected improvement in a completed task exceeds the call’s price and disclosure risk. The relevant unit is no longer cost per token. It is the incremental cost per accepted result.

Privacy narrows the route without closing it

Local execution creates a useful data boundary. Article 5 of the GDPR requires personal data to be adequate, relevant and limited to what is necessary. The UK Information Commissioner’s Office applies the same data-minimisation logic to AI systems and asks organisations to assess both security risks and the information genuinely required. Keeping source documents on the device can reduce the material disclosed to another processor and make a narrow cloud escalation easier to justify.

Cloud privacy has also improved, which limits the case for total localisation. OpenAI, Anthropic, Google Cloud and AWS state that commercial or API inputs are not used for model training by default. All four offer retention controls, and eligible customers can obtain zero-data-retention arrangements on parts of their platforms. Data residency, private networking, contractual terms and managed guardrails can make a hosted model acceptable for workloads that would once have required an internal deployment.

The remaining distinction concerns authority as much as storage. A remote model that receives a redacted question and returns text has a smaller action surface than a local agent with access to files, credentials and a shell. Perplexity’s own Numbat security work describes client agents crossing intended boundaries without a malicious prompt, while its SPACE architecture uses virtual machines, controlled egress and external credential stores to constrain long-running work. OWASP and the UK National Cyber Security Centre likewise treat prompt injection, excessive agency, supply-chain exposure and monitoring as system risks.

Locality moves responsibility rather than cancelling it. The sensible hybrid keeps data and execution local when exposure is material, then uses a cloud model through a narrow, logged and revocable interface.

The strongest case against the thesis

Open-weight progress could make frontier escalation unnecessary faster than cloud providers can defend their premium. Hugging Face counted more than 151,000 Qwen derivatives in its summer 2026 review, and it found permissive licences across much of the Chinese open-model ecosystem. Its integration of llama.cpp and GGML also shortens the path from a model repository to efficient local execution. Better post-training can improve a checkpoint without increasing the memory needed to run it.

That case is credible, but agent evidence still shows a difficult tail. MyPCBench places 184 tasks across a desktop containing seventeen simulated applications and 42,000 linked records. Its strongest tested agent completed 58.2% of tasks perfectly under the current 100-step evaluation and only 36% of tasks spanning at least seven applications. Several smaller models recorded no perfect completions on that cross-application slice. Perplexity’s own Terminal-Bench comparison leaves a 23-point gap between its local system and the frontier-only result before escalation.

Cloud prices are falling at the same time. Providers batch requests, cache repeated context and spread fixed infrastructure across many customers. One 2026 case study of an enterprise coding agent found that a 99.3% prompt-cache hit rate reduced realised API cost by 88.6%, below the amortised unit cost of the shared on-premise allocation in that deployment. Results from one organisation cannot settle the market, but they show why hardware ownership is not automatically cheaper.

The thesis fails when a local model meets the required quality on nearly every permitted task, routing errors disappear and frontier advice stops improving accepted outcomes enough to cover its price. Until then, local model portfolios and cloud models are complements with a contested boundary.

Buy escalation, not default access

An enterprise evaluating Portable Computer should replay its own work rather than compare model leaderboards. The test set should include private document processing, routine tool use, fresh research, long coding trajectories and cases where a wrong answer creates review or remediation work. Each run should record local utilisation, time to completion, disclosure at the cloud boundary, escalation frequency, frontier cost, retries and human acceptance.

Those measures expose the two failure modes. A machine that sits idle will not beat a cheap API on cost. A local model that produces plausible but rejected work will not beat it on productivity. Conversely, a frontier model called for every intermediate step turns hybrid architecture into an expensive relay. The target is a stable escalation frontier at which more local work lowers total cost without increasing failed outcomes.

This also changes cloud procurement. Buyers need a portfolio of approved models, clear data-handling routes and prices tied to completed work. Providers that can demonstrate a measurable improvement on the hard tail can preserve premium pricing even as their share of tokens falls. Providers that sell undifferentiated access will face the same compression already visible in routine inference.

The earlier argument that self-hosted agents require an operating model applies with more force once cloud escalation enters the loop. Ownership of the machine is only the first control. The durable advantage comes from knowing which work must remain inside it, which uncertainty deserves outside intelligence, and how much an accepted answer from the cloud is worth.

Sources

Portable Computer and the local stack

  1. Perplexity, "Introducing Portable Computer for local-first AI" https://www.perplexity.ai/ml/hub/blog/introducing-portable-computer-for-local-first-ai

  2. Perplexity Research, "A Local-First Agent for Private and Cost-Effective Knowledge Work" https://www.perplexity.ai/hub/blog/a-local-first-agent-for-private-and-cost-effective-knowledge-work

  3. Perplexity, "What is Personal Computer?" https://www.perplexity.ai/help-center/en/articles/14659663-what-is-personal-computer

  4. Perplexity Research, "Making SPACE: Secure and Efficient Runtimes for Long-Running Agents" https://www.perplexity.ai/hub/blog/making-space-secure-and-efficient-runtimes-for-long-running-agents

  5. Perplexity Research, "Securing Agents Across Perplexity’s Client Endpoints with Numbat" https://www.perplexity.ai/hub/blog/securing-agents-across-perplexity-s-client-endpoints-with-numbat

  6. NVIDIA, "NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents" https://blogs.nvidia.com/blog/local-ai-open-source-models-agents-nemotron/

  7. NVIDIA, "Personal AI Supercomputer Powered by Blackwell: NVIDIA DGX Spark" https://www.nvidia.com/en-us/products/workstations/dgx-spark/

  8. NVIDIA, "DGX Spark User Guide" https://docs.nvidia.com/dgx/dgx-spark/index.html

  9. NVIDIA, "Regulatory Compliance Information: DGX Spark" https://docs.nvidia.com/dgx/dgx-spark/compliance.html

  10. Qwen, "Qwen3.8-27B Model Card" https://huggingface.co/Qwen/Qwen3.8-27B

  11. Hugging Face, "State of Open Models: Summer 2026 Observations" https://huggingface.co/blog/state-of-open-models-summer-2026

  12. Hugging Face, "Use AI Models Locally" https://huggingface.co/docs/hub/en/local-apps

  13. Hugging Face, "GGML and llama.cpp Join HF to Ensure the Long-Term Progress of Local AI" https://huggingface.co/blog/ggml-joins-hf

  14. Hugging Face, "Hugging Face Response to the NTIA Request for Comment on Open-Weight Models" https://huggingface.co/api/resolve-cache/datasets/huggingface/policy-docs/6bf9c5b16bbe4cb34acb6f46a76384ec945f4f6e/2024_NTIA_Response.pdf

Production demand and inference economics

  1. Amelia Charles and Harpreet Arora, "DeepSeek Overtakes Google on Volume, Cost per Token Falls 13.6%" https://vercel.com/blog/deepseek-overtakes-google-on-volume-cost-per-token-falls

  2. Amelia Charles and Harpreet Arora, "Open-Weight Models Surge to 29% of Volume, Price per Token Flattens" https://vercel.com/blog/ai-gateway-production-index-july-2026

  3. Harpreet Arora and Yvonne Zhou, "AI Gateway Production Index" https://vercel.com/blog/ai-gateway-production-index

  4. Vercel, "Six LLM Routing Strategies for Teams Running Multi-Model Production Traffic" https://vercel.com/i/llm-routing-strategies

  5. OpenAI, "Models" https://developers.openai.com/api/docs/models

  6. OpenAI, "Advancing the Price-Performance Frontier with GPT-5.6" https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/

  7. Anthropic, "Models Overview" https://platform.claude.com/docs/en/models/overview

  8. Anthropic, "Pricing" https://platform.claude.com/docs/en/about-claude/pricing

  9. Google Cloud, "Vertex AI Pricing" https://cloud.google.com/vertex-ai/generative-ai/pricing

  10. Amazon Web Services, "Amazon Bedrock Pricing" https://aws.amazon.com/bedrock/pricing/

  11. Ba’Carri Johnson et al., "Introducing Granular Cost Attribution for Amazon Bedrock" https://aws.amazon.com/blogs/machine-learning/introducing-granular-cost-attribution-for-amazon-bedrock/

  12. Ege Erdil, "Inference Economics of Language Models" https://epoch.ai/publications/inference-economics-of-language-models

  13. Stanford Institute for Human-Centered Artificial Intelligence, "Artificial Intelligence Index Report 2026" https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf

  14. Hans Gundlach et al., "The Price of Progress: Price Performance and the Future of AI" https://arxiv.org/abs/2511.23455

  15. Guanzhong Pan et al., "A Cost-Benefit Analysis of On-Premise Large Language Model Deployment: Breaking Even with Commercial LLM Services" https://arxiv.org/abs/2509.18101

  16. Jonathan Knoop and Hendrik Holtmann, "Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs" https://arxiv.org/abs/2601.09527

  17. Sheng-Wei Peng, Yi-Hsun Lin and Yi-Pei Lee, "Inference Economics of Enterprise Coding Agents: A Case Study of Cloud vs. On-Premise LLMs" https://arxiv.org/abs/2607.13080

  18. Lenovo Press, "On-Premise vs Cloud: Generative AI Total Cost of Ownership, 2025 Edition" https://lenovopress.lenovo.com/lp2225-on-premise-vs-cloud-generative-ai-total-cost-of-ownership-2025-edition

Routing and agent evaluation

  1. Isaac Ong et al., "RouteLLM: Learning to Route LLMs from Preference Data" https://proceedings.iclr.cc/paper_files/paper/2025/hash/5503a7c69d48a2f86fc00b3dc09de686-Abstract-Conference.html

  2. Hao Li et al., "LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing" https://arxiv.org/abs/2601.07206

  3. Qitian Jason Hu et al., "RouterBench: A Benchmark for Multi-LLM Routing System" https://arxiv.org/abs/2403.12031

  4. Lingjiao Chen, Matei Zaharia and James Zou, "FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance" https://arxiv.org/abs/2305.05176

  5. Yu-Neng Chuang et al., "Learning to Route LLMs with Confidence Tokens" https://proceedings.mlr.press/v267/chuang25b.html

  6. Anshpreet Singh Bindra et al., "A Hybrid Local-Cloud Large Language Model Architecture for Privacy Preservation and Cost-Efficient Inference" https://doi.org/10.23919/INDIACom70271.2026.11525699

  7. Yasmin Moslem et al., "Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving" https://arxiv.org/abs/2606.27457

  8. Microsoft, "How Model Router Works in Microsoft Foundry" https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router-how-it-works

  9. Jason Wei et al., "BrowseComp: A Benchmark for Browsing Agents" https://openai.com/index/browsecomp/

  10. Terminal-Bench, "Terminal-Bench 2.1" https://www.tbench.ai/news/terminal-bench-2-1

  11. Lawrence Keunho Jang et al., "MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents" https://mypcbench.com/

Privacy, governance and security

  1. European Union, "Regulation (EU) 2016/679: General Data Protection Regulation" https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32016R0679

  2. Information Commissioner’s Office, "How Should We Assess Security and Data Minimisation in AI?" https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/how-should-we-assess-security-and-data-minimisation-in-ai/

  3. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile" https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence

  4. UK National Cyber Security Centre, "Guidelines for Secure AI System Development" https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development

  5. OWASP GenAI Security Project, "2025 Top 10 Risk and Mitigations for LLMs and Gen AI Apps" https://genai.owasp.org/llm-top-10/

  6. OWASP GenAI Security Project, "Agentic AI: Threats and Mitigations" https://genai.owasp.org/resource/agentic-ai-threats-and-mitigations/

  7. OpenAI, "Enterprise Privacy at OpenAI" https://openai.com/enterprise-privacy/

  8. OpenAI, "Offering Zero Data Retention for Frontier Models" https://openai.com/index/offering-zero-data-retention-for-frontier-models/

  9. Anthropic, "How Long Do You Store My Organization’s Data?" https://privacy.claude.com/en/articles/7996866-how-long-do-you-store-my-organization-s-data

  10. Google Cloud, "Vertex AI and Zero Data Retention" https://docs.cloud.google.com/vertex-ai/generative-ai/docs/vertex-ai-zero-data-retention

  11. Amazon Web Services, "Amazon Bedrock Security, Privacy, and Responsible AI" https://aws.amazon.com/bedrock/security-privacy-responsible-ai/