"NUNTII EX MACHINA"
DOCUMENTING THE RACE TO AGI

THE
AGENTIC
TIMES_

▮ NEWS_TICKER.LIVE REC
▸ Ryan Serhant Reveals How He Turns AI Agents Into Million-Dollar Sales Forbes · 21.07.2026▸ Chinese Tech Firms Pitch AI Agents as the Future of Smartphones citynewsservice.cn · 21.07.2026▸ Powerful AI Models Happily Being Given Away for Free RealClearMarkets · 21.07.2026▸ GitLab 19.2 Puts AI Agents to Work on the Security Backlog infoq.com · 21.07.2026▸ OpenAI’s GPT-Red: How AI Models Are Now Training Each Other to Be More Secure quasa.io · 21.07.2026▸ MCP update prepares AI agents for widespread deployment Techzine Global · 21.07.2026▸ New ENCFORGE Ransomware Targets AI Model Files in Langflow RCE Attack The Hacker News · 21.07.2026▸ In Depth: Chinese Smartphone-Makers Bet on AI Agents Caixin Global · 21.07.2026▸ WebMCP: Bridging the Gap Between AI Agents and the Web devmio · 21.07.2026▸ Boomi study finds AI agent trust lags enterprise adoption InsiderPH · 21.07.2026▸ Rapidus and Cadence partner to advance AI-driven SoC design New Electronics · 21.07.2026▸ Agentic AI and Smart Data: The Architecture of the UK’s New Open Finance Framework… Finextra Research · 21.07.2026▸ Databricks: US$188bn Valuation, Genie One and Agentic AI AI Magazine · 21.07.2026▸ Nvidia targets simulation bottlenecks with AI agent expansion InsiderPH · 21.07.2026▸ China Weighs Export Controls on AI Models, Including Open Weight LLMs trendingtopics.eu · 21.07.2026▸ Research says JadePuffer Ransomware wipes off data on AI Model Infrastructure Cybersecurity Insiders · 21.07.2026▸ Reported US push to ban Chinese AI models reflects anxiety over eroding tech hegemony… Global Times · 21.07.2026▸ Webinar: Can AI agents finally automate data testing? QA Financial · 21.07.2026
▮ MODEL_FEED.LIVE REC
▸ Nemotron-Labs-Audex-30B-A3B NVIDIA · 30B · MoE · 06.07.2026▸ Nemotron-Labs-Audex-2B NVIDIA · 2B · 06.07.2026▸ DeepSeek-V4-Flash-DSpark DeepSeek · Open language model · 27.06.2026▸ Qwen-AgentWorld-35B-A3B Qwen · 35B · MoE · multimodal · 22.06.2026▸ GLM-5.2 Zhipu · Open language model · 16.06.2026▸ North-Mini-Code-1.0 Cohere · coding · 05.06.2026▸ DeepSeek-V4-Pro DeepSeek · Open language model · 22.04.2026▸ DeepSeek-V4-Flash DeepSeek · Open language model · 22.04.2026▸ granite-4.1-8b IBM · 8B · 06.04.2026▸ granite-4.1-3b IBM · 3B · 06.04.2026▸ Mamba2-primed-HQwen3-8B-Instruct Amazon · 8B · instruct · 31.03.2026▸ Falcon-OCR TII · Open language model · 22.02.2026▸ tiny-aya-base Cohere · Open language model · 13.02.2026▸ tiny-aya-global Cohere · Open language model · 13.02.2026▸ GLM-4.7-Flash Zhipu · Open language model · 19.01.2026▸ Falcon-H1R-7B TII · 7B · 29.10.2025
▮ PERSPECTIVE / 31.03.2026 · 4 MIN READ

AI Factories: It's Basic Tokenomics

AI Factories: It's Basic Tokenomics

The reality of the AI Factory is here. Data centers are no longer just places where compute lives, they’re production facilities. And the unit of output isn’t a calculation or a query. It’s intelligence. A token.

From where I sit, working with some of the companies building this infrastructure, that framing changes everything about how enterprises should be thinking about AI investment. Because if AI is a manufacturing operation, then the questions that matter aren’t just “which model?” or “which cloud?” They’re: what does it cost to produce a token? Who produces them most efficiently? And what happens when the factory runs out of power?

The metric your AI budget is missing

Most enterprises buying AI at scale today are making decisions based on headline model benchmarks and sales conversations. The challenge isn’t awareness, CTOs know scaling AI isn’t cheap. Fewer have the internal benchmarking maturity to drive that cost down and measure return against it.

That matters because not all token generation is equal. Inference workloads (the live, real-time production of AI responses) have a completely different cost profile to training runs. Throughput, latency, and GPU utilisation rates interact in ways that can mean the difference between an AI deployment that scales economically and one that quietly bleeds budget.

The companies getting this right are building internal benchmarking capability. They’re asking providers not just what the model can do, but how the infrastructure behind it is optimised, batching strategies, quantisation approaches, and the data layer performance that determines how fast tokens actually flow.

A market that’s finally more diverse

The hyperscalers like AWS, Azure, Google Cloud typically dominate the conversation, and for good reason. Their integration depth, ecosystem breadth, and enterprise relationships are genuinely hard to replicate. AWS’s investment in custom silicon like Trainium and Inferentia is also a serious play at owning the cost curve from the chip up.

But the challenger layer is more credible than it’s often given credit for. Specialist GPU cloud providers (CoreWeave, Nebius, Nscale, Lambda) and sovereign infrastructure players are competing hard on price per token and deployment flexibility, and winning massive deals with enterprises who need more than a hyperscaler’s standard menu. Meanwhile, the storage layer, long underestimated in AI deployments, is now recognised as a critical performance variable. Purpose-built AI storage solutions are increasingly part of what separates an efficient token pipeline from an expensive one. The intelligent procurement decision right now isn’t picking a winner. It’s building a portfolio.

The power to succeed

There’s one dimension of tokenomics that enterprises are only beginning to factor in: power. Running AI Factories at scale is extraordinarily energy-intensive, and as data center capacity tightens across key markets, power availability is becoming a genuine constraint on some AI ambitions.

The good news is that infrastructure efficiency and sustainability are increasingly the same conversation. Smaller, distilled, and quantised models (reducing memory use and speed up inference) can deliver capability at a fraction of the energy cost of frontier models. Hardware-software co-optimisation (matching workloads precisely to the right compute) is also reducing wastage. The providers investing in this aren’t just being responsible. They’re building a structural cost advantage that will only become more and more critical.

Bending the curve

The AI infrastructure market is moving faster than most enterprise roadmaps. The cost-per-token curve is compressing too. The competitive landscape is shifting on what seems like a monthly basis, and the energy question is becoming a board-level issue.

My suggestion, treat tokenomics with the same rigour you’d apply to any other unit of production cost. Understand what you’re buying, how it’s made, and what the real price of scaling that effort it looks like. The companies that win the next phase of AI won’t necessarily have the best models. They’ll have the most efficient factories for AI.

END OF TRANSMISSION ▮ ◂ MORE PERSPECTIVE