"NUNTII EX MACHINA"
DOCUMENTING THE RACE TO AGI

THE
AGENTIC
TIMES_

▮ NEWS_TICKER.LIVE REC
▸ Ryan Serhant Reveals How He Turns AI Agents Into Million-Dollar Sales Forbes · 21.07.2026▸ Chinese Tech Firms Pitch AI Agents as the Future of Smartphones citynewsservice.cn · 21.07.2026▸ Powerful AI Models Happily Being Given Away for Free RealClearMarkets · 21.07.2026▸ GitLab 19.2 Puts AI Agents to Work on the Security Backlog infoq.com · 21.07.2026▸ OpenAI’s GPT-Red: How AI Models Are Now Training Each Other to Be More Secure quasa.io · 21.07.2026▸ MCP update prepares AI agents for widespread deployment Techzine Global · 21.07.2026▸ New ENCFORGE Ransomware Targets AI Model Files in Langflow RCE Attack The Hacker News · 21.07.2026▸ In Depth: Chinese Smartphone-Makers Bet on AI Agents Caixin Global · 21.07.2026▸ WebMCP: Bridging the Gap Between AI Agents and the Web devmio · 21.07.2026▸ Boomi study finds AI agent trust lags enterprise adoption InsiderPH · 21.07.2026▸ Rapidus and Cadence partner to advance AI-driven SoC design New Electronics · 21.07.2026▸ Agentic AI and Smart Data: The Architecture of the UK’s New Open Finance Framework… Finextra Research · 21.07.2026▸ Databricks: US$188bn Valuation, Genie One and Agentic AI AI Magazine · 21.07.2026▸ Nvidia targets simulation bottlenecks with AI agent expansion InsiderPH · 21.07.2026▸ China Weighs Export Controls on AI Models, Including Open Weight LLMs trendingtopics.eu · 21.07.2026▸ Research says JadePuffer Ransomware wipes off data on AI Model Infrastructure Cybersecurity Insiders · 21.07.2026▸ Reported US push to ban Chinese AI models reflects anxiety over eroding tech hegemony… Global Times · 21.07.2026▸ Webinar: Can AI agents finally automate data testing? QA Financial · 21.07.2026
▮ MODEL_FEED.LIVE REC
▸ Nemotron-Labs-Audex-30B-A3B NVIDIA · 30B · MoE · 06.07.2026▸ Nemotron-Labs-Audex-2B NVIDIA · 2B · 06.07.2026▸ DeepSeek-V4-Flash-DSpark DeepSeek · Open language model · 27.06.2026▸ Qwen-AgentWorld-35B-A3B Qwen · 35B · MoE · multimodal · 22.06.2026▸ GLM-5.2 Zhipu · Open language model · 16.06.2026▸ North-Mini-Code-1.0 Cohere · coding · 05.06.2026▸ DeepSeek-V4-Pro DeepSeek · Open language model · 22.04.2026▸ DeepSeek-V4-Flash DeepSeek · Open language model · 22.04.2026▸ granite-4.1-8b IBM · 8B · 06.04.2026▸ granite-4.1-3b IBM · 3B · 06.04.2026▸ Mamba2-primed-HQwen3-8B-Instruct Amazon · 8B · instruct · 31.03.2026▸ Falcon-OCR TII · Open language model · 22.02.2026▸ tiny-aya-base Cohere · Open language model · 13.02.2026▸ tiny-aya-global Cohere · Open language model · 13.02.2026▸ GLM-4.7-Flash Zhipu · Open language model · 19.01.2026▸ Falcon-H1R-7B TII · 7B · 29.10.2025
▮ DISPATCHES / 22.04.2026 · 3 MIN READ

Agentic AI Matches Human Economists on Accuracy — Beats Them on Consistency

Agentic AI Matches Human Economists on Accuracy — Beats Them on Consistency

A well-established problem is that the same data, given to different research teams with the same question, produces very different answers. A landmark 2025 study by Huntington-Klein et al. demonstrated this at scale: 146 economist teams were each asked to estimate the employment effects of the DACA immigration policy using the same dataset. The spread of answers was striking with estimates ranged across a wide band, driven by thousands of small analytical choices each team made independently about sample construction, research design, and statistical method.

A new working paper from the Federal Reserve Board reran that same exercise with three agentic AIs: Codex with GPT-5.4, Codex with GPT-5.3-Codex, and Claude Code with Opus 4.6. Each model ran 100 independent instances across three progressively constrained versions of the task — 900 total runs — each starting from scratch with no knowledge of the others.

Across all three tasks, the tendency of AI estimates tracked closely with human results. The more significant finding though was in dispersion (the spread, variability, or scattering of data points around a central value). Human economists produced estimates with considerably wider tails: standard deviations up to nine times larger than Codex GPT-5.4 in Task 1, and ranges extending into extreme outliers that no AI model approached. Prescribing the research design in Task 2 reduced AI dispersion substantially — standard deviations fell by 20–45% across models — while human dispersion barely changed. Essentially, AI systems are more sensitive to structured constraints, and when given them, converge more tightly on defensible answers.

The second half of the paper is more striking. Author Serafin Grundl formed 300 comparison groups, each containing one submission from each AI system and one from a human economist. Multiple AI reviewer models then independently ranked each group on methodological quality, with code-level evidence required for every substantive claim. The ranking was identical regardless of which model did the reviewing: GPT-5.4 first, GPT-5.3-Codex second, Opus 4.6 third, humans fourth. Crucially, models showed almost no favouritism toward their own submissions. Opus 4.6 consistently ranked GPT-5.4 and GPT-5.3-Codex ahead of itself!

The paper does not declare victory as such. AI models make errors and cannot yet be assumed error-free simply because recent models improved on their predecessors. The author notes that both caveats apply to human researchers too, and that running multiple AI instances makes it easier to detect errors and map the space of reasonable analytical choices than the equivalent human exercise.

The conclusion is measured but pointed. Agentic AI systems can now perform substantive economics research at human-comparable quality and can do so at a scale and consistency that human teams cannot seem to match.

END OF TRANSMISSION ▮ ◂ MORE DISPATCHES