InternationalItaliano中文
IL PROGRESSO

Independent journalism on global markets, technology, and the forces reshaping the world economy

Ufficio Emissioni · VeneziaEmissione N. 1412
Home /Technology /Emissione
Technology01 MIN

OpenAI’s Opaque Reasoning: New Astra Technique Tests AI Safety Norms

OpenAI is preparing to release its Astra model with a reasoning technique known as “recurrent depth,” a method that lets the model operate outside the sequential thinking that characterizes most current reasoning systems. The technique, als…

OpenAI’s Opaque Reasoning: New Astra Technique Tests AI Safety Norms

OpenAI is preparing to release its Astra model with a reasoning technique known as “recurrent depth,” a method that lets the model operate outside the sequential thinking that characterizes most current reasoning systems. The technique, also called “opaque recurrence,” will likely make Astra’s internal chain of thought far more difficult for outsiders to monitor, and its emergence has prompted sharp warnings from AI safety researchers who say the practice threatens the transparency that underpins current oversight efforts.

Chain-of-thought reasoning is the mechanism by which a model articulates its steps toward a conclusion in a human-readable sequence. That transparency has become a foundational safety tool: researchers can inspect how a model reaches a decision, detect biases or errors, and intervene before deployment. Recurrent depth undermines that arrangement by allowing the model to revisit and refine its internal computational states in ways that do not map cleanly onto a linear trace. The result is a model that may reason more efficiently, but whose decision-making is significantly harder to audit.

Astra’s use of the technique is reportedly limited, but the reaction among safety specialists has been swift. Buck Shlegeris, chief executive of the alignment research group Redwood, said he was “extremely concerned” by the reporting, while Redwood’s chief scientist, Ryan Greenblatt, warned that opaque reasoning could scale faster than conventional chain-of-thought approaches and eventually remove all reasoning from visible channels. Greenblatt said he hoped it was not too late to avoid the most concerning architectures. Zvi Mowshowitz, a longtime AI safety advocate, went further, suggesting that legislation may be needed to prevent a “race to the bottom” among AI labs, arguing that more intensive use of the technique would likely damage monitorability.

The tension here is between capability and oversight, a trade-off that has defined the frontier of AI development. Opaque recurrence may offer genuine performance advantages, but it erodes the very transparency that safety teams have relied on to catch harmful behavior before systems reach the public. OpenAI’s chief scientist, Jakub Pachocki, defended the company’s position, stating that preserving and utilizing chain-of-thought monitoring has been a core goal of its research program since its first reasoning models. The company’s framing suggests it views the Astra deployment as a contained experiment rather than an abandonment of transparent reasoning. But critics contend that such boundaries are difficult to hold: once a technique proves valuable at a limited scale, competitive pressure to expand it grows considerably, especially in a market where labs are racing for capability advantages.

The wider implication is that AI governance may need to shift from relying on voluntary transparency to establishing enforceable standards. If the industry’s leading labs treat monitorability as a negotiable feature rather than a fixed requirement, the tools that allow society to assess AI risk begin to disappear. That prospect is what makes the Astra decision significant not merely as a technical choice, but as a signal about the future direction of frontier model development. The question is whether OpenAI and its peers will treat chain-of-thought fidelity as a baseline commitment or as a constraint to be optimized away as capabilities advance. The answer will shape not only the safety of individual models, but the broader credibility of the industry’s promises about responsible deployment.

Source & Credits

Originally reported by Slashdot.

Written for Il Progresso by Zhicheng Wang.

↑ Torna alla prima pagina