• The Prohuman
  • Posts
  • AI’s math progress is creating a new problem

AI’s math progress is creating a new problem

Plus: Grok 4.7 starts at $2 per million tokens

In partnership with

Hello, Prohuman

Today, we will talk about these stories:

  • OpenAI is asking mathematicians for oversight

  • Grok 4.7 gets cheaper and stronger

  • NVIDIA says task completion is the real AI test

Elon's new company is private. These 3 tickers aren't.

The next Apple may already exist. Insider sources say Elon has spent two years building a secret device inside Tesla's facilities — one he claims will be "10x bigger than the largest product in history."

There's just one problem: the company is private, and unless you know Elon personally, you can't buy a single share. That was true until Guardian's research team found three public ticker symbols sitting in the launch supply chain.

Click here to see all 3 tickers, free of charge.

You won't hear these names on CNBC — Wall Street hasn't published a word on the connection. But when the launch hits September 21, that quiet ends.

Some are already calling this the biggest opportunity since AI. For anyone who missed Apple before the iPhone, this may be a second look at that kind of setup.

OpenAI’s math model is moving faster than expected

Image credits: OpenAI

OpenAI says an internal model has resolved more than 100 long-standing open math problems since training began August 28.

That includes its claimed resolution of the Navier-Stokes Millennium Prize problem, and OpenAI says the pace surprised its own mathematicians. It has now started working with an independent advisory group that includes Timothy Gowers, Martin Hairer, Edward Witten and six other mathematicians.

The interesting part is OpenAI admitting that publishing every new result immediately may create problems for the field. When a model can work through open problems this quickly, academic credit, verification and decisions about releasing results become practical issues researchers have to handle at their desks.

The group can publicly criticize OpenAI and will receive no payment from the company. But it will not advise OpenAI on slowing its internal math research.

That boundary may matter most.

 Grok 4.7 is competing on price now

Image credits: x.AI

Grok 4.7 costs $2 per million input tokens and $6 per million output tokens.

SpaceXAI says its larger new model was trained longer on difficult tasks that can take hours to complete. On CursorBench 4.0, Grok 4.7 scored 46.3%, ahead of the 41.7% figure SpaceXAI reports for GPT-5.6 Sol Max, while Fable 5.1 Max scored 51.8%.

The pricing stands out. Developers staring at API bills now have another capable model priced aggressively enough to test on real workloads instead of benchmarks alone.

SpaceXAI is also clearly pushing beyond chat and into coding, documents, presentations, legal work and other tasks where models need to keep working without constant supervision.

Grok 4.7 is already available through Cursor, Grok Build, its API and third-party coding tools. That makes switching or running side-by-side tests relatively easy.

The next useful numbers will come from actual developer bills.

Agent benchmarks are starting to measure finished work

Image credits: NVIDIA

A correct tool call means very little if the job still fails.

NVIDIA argues that agent evaluation should track the final state of a task, such as whether a refund actually posted or a software bug was fixed. That requires running agents inside executable environments where every tool call changes something and can be checked afterward.

This feels much closer to how companies should evaluate agents. A team does not care that an agent correctly called an API twelve times if someone still has to open the laptop and finish the task manually.

The useful metrics become task success, consistency and cost per successful job, with individual tool calls kept mainly for diagnosing failures.

NVIDIA says its Nemotron 3.5 Lightning reaches 86% accuracy on PinchBench while completing 10,000 tasks 30% faster than Qwen3.6 35B at comparable accuracy.

The harder question is what percentage of your actual work finishes correctly.

Prohuman team

Covers emerging technology, AI models, and the people building the next layer of the internet.

Founder

Writes about how new interfaces, reasoning models, and automation are reshaping human work.

Founder

Free Guides

Explore our free guides and products to get into AI and master it.

All of them are free to access and would stay free for you.

Feeling generous?

You know someone who loves breakthroughs as much as you do.

Share The Prohuman it’s how smart people stay one update ahead.