• The Prohuman
  • Posts
  • Google's new Gemini models focus on efficiency

Google's new Gemini models focus on efficiency

Plus: NVIDIA's AI push now runs through Texas

In partnership with

Hello, Prohuman

Today, we will talk about these stories:

  • Gemini 3.6 Flash cuts costs and speeds up AI agents

  • The AI race is moving onto factory floors

  • Microsoft and Mistral deepen AI partnership

You've seen the AI demos. Viktor does it without you watching.

The AI tool you tried last quarter waited for a prompt, hallucinated a number, then asked if you'd like a summary.

Viktor opened a PR at 2am, rebased it against main, ran your test suite, and posted a note in #eng: "Two flaky tests in payments service, both pre-existing. Recommended merging after fixing them." Then drafted the customer reply for the support ticket the bug created.

That's 619K autonomous actions per day across 20,000+ teams. Not chat replies. Real work shipped to GitHub, Stripe, Linear, Notion, and 3,000+ other tools, from inside Slack and Microsoft Teams.

You don't supervise him any more than you supervise a senior engineer.

SOC 2 certified. Your data never trains models.

"It's what you probably originally thought AI was going to be when you first heard of it in sci-fi movies." Tyler, CEO.

Efficiency is becoming the real AI race

Image Credits: Google

Google's biggest claim isn't that Gemini 3.6 Flash is smarter. It's that it gets more done while using fewer tokens.

The company introduced Gemini 3.6 Flash, 3.5 Flash-Lite, and a specialized cybersecurity model called 3.5 Flash Cyber. The headline numbers are practical: 3.6 Flash uses 17% fewer output tokens than its predecessor, while Flash-Lite delivers up to 350 output tokens per second at a much lower price. Google also confirmed it has started pre-training Gemini 4.

That tells me the market is shifting. Developers already have strong models to choose from, so cost, latency, and reliability are becoming the deciding factors for production AI systems. You notice this when the work has to run all day without slowing down or becoming expensive.

The cybersecurity model is also worth watching. Google is limiting access through CodeMender and trusted partners, which suggests the company sees security models as tools that need tighter controls than general-purpose AI.

The next question is whether developers value lower operating costs more than another jump in benchmark scores.

AI control is becoming a selling point

Image Credits: Microsoft

This deal isn't about adding another AI model. It's about giving large organizations more control over where AI runs.

Microsoft and Mistral are expanding their partnership with a multibillion-dollar agreement that adds more Europe-based GPU capacity while bringing Mistral's latest models into Microsoft Foundry and Copilot Studio. The companies are also making it possible to deploy those models across public cloud, cloud-connected systems, or fully disconnected environments for regulated industries.

The timing makes sense. Governments, healthcare providers, manufacturers, and financial firms increasingly want frontier AI without giving up control over sensitive data or depending entirely on public cloud infrastructure. That's becoming a competitive feature instead of a niche requirement. Someone managing critical systems cares less about the newest benchmark and more about keeping operations running under strict rules.

This partnership also shows Microsoft's strategy is widening. Instead of pushing only its own models, it's positioning Azure as the place where customers can choose the model that best fits their regulatory and operational needs.

The next phase of AI competition may depend as much on deployment flexibility as model performance.

AI infrastructure needs factories, not just chips

Image Credits: NVIDIA

The biggest announcement isn't another AI chip. It's a new factory built to make them at scale.

NVIDIA and manufacturing partner Wistron opened a 324,000-square-foot facility in Fort Worth, Texas, where production has started on the GB300 Grace Blackwell Ultra Superchip and will soon expand to the Vera Rubin Superchip. The company says the site is part of a broader $700 million investment, with more than 500 jobs already created and plans to reach 1,000 by the end of the year.

This feels like a shift in where the AI competition is heading. The conversation is moving beyond model benchmarks toward manufacturing capacity, supply chains, and the ability to produce advanced systems in large volumes. You can hear the machines running long before you see the finished hardware.

The factory was also designed as a digital twin before construction began, showing that AI is now helping build the infrastructure that powers the next generation of AI itself.

The next advantage may belong to the companies that can manufacture as quickly as they innovate.

Prohuman team

Covers emerging technology, AI models, and the people building the next layer of the internet.

Founder

Writes about how new interfaces, reasoning models, and automation are reshaping human work.

Founder

Free Guides

Explore our free guides and products to get into AI and master it.

All of them are free to access and would stay free for you.

Feeling generous?

You know someone who loves breakthroughs as much as you do.

Share The Prohuman it’s how smart people stay one update ahead.