There is an elegant irony in the fact that OpenAI chose the name "Astra" for its next major model. In Greek tradition, Astra was the Titaness associated with the light of the stars and, just as tellingly, with justice and the precision of what cannot be avoided. It is precisely that notion of precision the company is trying to sell now, announcing that Astra has solved — or at least decisively advanced — ten long-standing mathematical problems using formal verification with the Lean assistant.
Mathematics, as any student learns early, rarely bends to shortcuts. It demands proof, not intuition. And that is exactly where Astra's conceptual turn lies. Instead of merely generating plausible text, the model can collaborate with itself across different parts of a larger problem, breaking apart what once looked like a monolithic block and assembling the solution piece by piece. Formal verification with Lean acts as a kind of merciless referee: every step must be confirmed by mechanical rules, leaving no room for the comfortable rhetoric that masks the wrong answers of ordinary language models. It is a qualitative leap over machines that merely memorize patterns and regurgitate responses with an air of authority, never able to justify their own reasoning.
The cost of that ambition is brutal. According to OpenAI, solving the ten challenges consumed roughly two thousand dollars in API tokens — a number that illustrates, in concrete terms, the energy and computational price of trying to turn intuition into proof. The proof files are public, a gesture of transparency that invites the community to audit the work. But the model itself remains private, which raises an inevitable question: public proof, closed intelligence.
Anthropic's reply came in the same competitive currency. The company says its public model, Claude Fable, independently solved five of the same challenges. None of this happens in a historical vacuum: OpenAI is coming off successive generations — Sol, Terra and Luna — that have progressively refined how these machines reason. It has not yet been decided whether Astra will ship as GPT-5.7, GPT-6, or under a completely new name, and that naming indecision almost reads like a confession that the company does not yet know how to position what it has built.
What does it mean, after all, for an AI to "prove mathematics" when it comes to credibility? For me, it is the most serious promise this industry has ever made. By betting on formal verification, OpenAI stops asking us to trust its models and instead asks us to trust the rules of logic — an intelligent transfer of trust, even if partial. The race between OpenAI and Anthropic is good for the field, because what is at stake is not a marketing benchmark but the ability to demonstrate reasoning in an auditable way. Competition, in this case, works as an accelerator of quality, forcing each lab to expose its methods with greater rigor. The question that remains, and that unsettles me, is simple: if ten mathematical problems cost two thousand dollars in tokens, how much will it cost to prove something genuinely important to everyday life — and who will pay that bill? The public needs to understand that formal verification, however powerful, still operates in a narrow and well-defined territory, far from the ambiguity of real life.
Sources: Moneycontrol, Implicator AI, Technoid
✓ Independent sources cross-checked and verified before publishing