Anthropic is automating AI research

Plus: Flight tracking moves inside Google AI Mode

In partnership with

Hello, Prohuman

Today, we will talk about these stories:

  • Anthropic tests AI improving AI

  • Google AI Mode is becoming a travel agent

  • AI security scanners disagree with themselves

Stop Paying for 10 Tools. One AI Does It All.

Most e-commerce sellers are running their store across 6 to 10 separate tools — and spending more time managing software than growing their business. StoreClaw replaces your entire stack with one autonomous AI engine that monitors competitors, optimizes listings, automates marketing, and tracks real profit across Shopify, Amazon, and beyond.

It doesn't wait for you to ask. It runs 24/7 in the background, so you wake up to a full dashboard instead of a list of things you forgot to check.

Connect your store, and StoreClaw gets to work — no prompts, no complex setup, no six-app stack.

Free to start. No credit card required.

Anthropic’s automated researcher beat experienced humans

Anthropic just put a price on automated AI research: roughly $4 an hour.

Its new system searches existing research, chooses a training method, runs it for 30 minutes, and keeps approaches that improve alignment benchmarks. Across 10 tests for unwanted model behavior, it improved every benchmark without hurting overall performance.

The striking part is speed. Anthropic says its best automated method beat proposals from experienced human researchers within six hours, while those researchers cost about $150 an hour.

I think the immediate implication is fairly practical: labs can run far more research experiments without adding people at the same rate. You can picture racks of machines running these short training cycles overnight while researchers decide which results deserve closer attention.

There is an important constraint. The system can only optimize against benchmarks humans have defined and maintained, so weak measurements could still produce convincing but misleading progress.

How much research work stays human once these systems get better?

Google is pulling travel booking into AI Mode

Google AI Mode can now track airfare, surface hotels, and help users move much closer to completing a booking.

That changes the product. Google says flight results pull current prices from more than 300 airlines and travel sites, with price tracking now available in over 180 countries.

Hotel booking goes further. Users can describe a trip, compare options and reviews, then continue through Google’s partners and pay with Google Pay after checking details like room type and cancellation policy.

I think travel is a strong test case for AI agents because the work is repetitive, specific, and often happens with a phone open beside a suitcase. Google already controls much of the discovery process, so extending that experience into booking could reduce the number of times users leave Search.

Travel sites should watch closely. Booking.com, Expedia, Marriott and others are participating today, but Google increasingly controls the interface where the customer makes the decision.

How much of the booking relationship will those partners keep?

AI security scanners still produce a lot of noise

AI AppSec tools have a consistency problem.

Contrast Security tested three scanners on the same codebase and found they agreed on just 5% of findings. One scanner repeated against identical code reproduced only 17% of its own results.

The cost gap is worse. Scanning 2 million lines of code cost about $315 in API charges, while triaging the resulting findings cost roughly $128,000.

That is the part I would pay attention to because security teams already have large queues, with the average monitored application carrying 106 vulnerability findings. Adding cheap detection that produces unstable results can easily create more work for the people deciding what gets fixed Monday morning.

Meanwhile, attackers keep moving quickly. Contrast recorded 42 confirmed viable exploit attempts per application each month, while the average highest-severity vulnerability took 92 days to remediate.

If AI keeps finding more issues, who decides which ones actually deserve attention?

Prohuman team

Covers emerging technology, AI models, and the people building the next layer of the internet.

Founder

Writes about how new interfaces, reasoning models, and automation are reshaping human work.

Founder

Free Guides

Explore our free guides and products to get into AI and master it.

All of them are free to access and would stay free for you.

Feeling generous?

You know someone who loves breakthroughs as much as you do.

Share The Prohuman it’s how smart people stay one update ahead.