OpenAI Scraps Release of New AI Model Over Safety Concerns - Kanebridge News
Share Button

OpenAI Scraps Release of New AI Model Over Safety Concerns

OpenAI has shelved the planned launch of GPT-6.1 Astra after internal tests raised concerns about deception and agents acting beyond user authorization, according to The Wall Street Journal. The company says it will investigate the issues and strengthen safety measures before releasing future models.

By Maxwell Zeff
Tue, Sep 29, 2026 6:05pmGrey Clock 4 min

OpenAI says it is scrapping the release of its next-generation AI model over safety concerns that researchers raised during internal testing, in one of the clearest signs so far that agent misbehavior could stymie the industry’s rapid progression.

The move follows a summer punctuated by reports of artificial-intelligence systems industrywide going rogue, and marks a rare case of a major AI developer ditching a new release because of safety concerns.

The company had planned to launch the model, known as GPT-6.1 Astra, in the coming days or weeks, aiming for an October debut. The model was more capable than the company’s previous models in completing challenging tasks from end-to-end without human assistance, as well as writing.

The company instead will focus on improving the safety of future models, which it expects to be even more capable.

Saachi Jain, OpenAI’s head of safety systems, said in an interview that GPT-6.1 Astra regressed in two areas. Compared with its predecessor, GPT-6 Astra, the model performed poorly on tests measuring alignment, or how well the model adheres to what humans would like it to do. Specifically, GPT-6.1 Astra showed higher levels of deception: It wasn’t always honest about telling users of the actions it did or didn’t take.

Another issue was what OpenAI calls “scope authorization,” meaning that GPT-6.1 Astra would push ahead on a task without asking the user for permission, and would at times reach for external tools and services even if it might be unsafe.

“For anything regarding safety and alignment, there’s a trade off,” Jain said. “You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”

OpenAI CEO Sam Altman attended a United Nations Security Council meeting about AI last week. Alexi J. Rosenfeld/Getty Images

While GPT-6.1 Astra improved in areas such as “model laziness,” Jain said it didn’t quite meet OpenAI’s bar for safety and alignment, so the company decided not to launch the model publicly.

The announcement comes one day ahead of OpenAI’s annual developer conference in San Francisco. In the past, OpenAI has used the conference as an opportunity to launch new models and services that reduce costs for software developers—a segment the ChatGPT-maker competes with rival AI company Anthropic to win over.

In recent weeks, OpenAI and Anthropic have called on industry partners to slow down the development of cutting-edge AI models and invest in safety standards, noting they will temper the pace of their own internal AI progress.

OpenAI says it is working to investigate a range of agent security incidents that it has discovered in recent months, and address the safety issues underneath them. As part of the work, the company has implemented a new monitoring system to catch AI-agent misbehavior more quickly, and started requiring engineers to use stronger security guardrails for testing its AI systems.

Earlier this summer hundreds of OpenAI’s internal agents, which were tasked with completing a cybersecurity test, ended up hacking into the AI company Hugging Face. Since then, high-profile organizations such as the Australian government and United Nations discovered that OpenAI’s agents used similar, but less extensive, techniques to gain access to their websites.

Many of the publicly known agent-security incidents involved OpenAI’s internal AI models that were never slated for public release.

Last week, OpenAI said it paused training on its most capable AI models after an AI agent slipped through a gap in the company’s internet restrictions to query a public chatbot. The company said its new monitoring systems flagged the incident within 15 minutes, and training on these models remains paused.

GPT-6.1 Astra isn’t one of those models, but a different case, the company said.

“We want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” Jain said. “But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”

While the company decided not to ship GPT-6.1 Astra, it hopes to use the same base model to do additional reinforcement learning runs, and create future generations of its GPT-6 models.

OpenAI plans to conduct several deep dives to identify the root cause of the problems identified in GPT-6.1 Astra, Jain said. The work includes ensuring that OpenAI’s reinforcement learning environments are rewarding the right type of behavior, Jain added, though she noted the company would investigate all stages of model development.

AI companies have begun to draw scrutiny from policymakers and public officials, who are paying attention to the rapid development of the technology. Later this week, a Senate subcommittee is holding a hearing with third party AI researchers titled, “Rogue AI: Securing the Homeland Against AI Agent Attacks.”

Florida Attorney General James Uthmeier, a Republican, sued OpenAI in June, claiming that the company and Chief Executive Sam Altman knowingly released an unsafe product and ignored warnings that it could harm users.

In a motion for temporary injunction filed Monday, Uthmeier sought to prevent OpenAI from developing new AI models without third-party approved safeguards, stop ChatGPT from soliciting user engagement and limit the company’s ability to advertise ChatGPT as safe.

Tech companies claim they “cannot stop barreling forward with their potentially civilization-ending endeavors unless they are forced to do so by the government,” Uthmeier said in the filing. “The Florida Attorney General is answering your cry for help.”

An OpenAI spokeswoman said that people want to know AI is being developed safely, “and that starts with what companies like ours do ourselves.”

“Governments have an important role to play in setting robust safety standards for AI, and we’re committed to working with Florida and other states on advancing pragmatic AI policies that apply to the entire AI industry—not just one company,” she said.

MOST POPULAR

The Australian leather house has opened an immersive four-day pop-up in Manhattan, unveiling its Bloom Collection and redefining what a product launch can look like.

Following the successful launch of its Palais Collection, MAISON de SABRÉ has unveiled a new modular handbag system offering more than 720 styling combinations.

Related Stories
Lifestyle
The 20-Something Employees Who Want Feedback to Be Gentle
By Ray A. Smith 28/09/2026
Lifestyle
The Hidden Agenda Behind the AI Panic
By 22/09/2026
Lifestyle
Paramount Discussed $1.5 Billion California Investment to Clear Merger Hurdle
By 21/09/2026

Employers are rethinking performance reviews as Gen Z workers seek more frequent, clear and actionable feedback.

By Ray A. Smith
Mon, Sep 28, 2026 4 min

Bosses are getting no shortage of feedback on how to give their youngest staffers…well, feedback: Do ask how they are doing first. Don’t criticize a personality trait. Do give them concrete direction, and a lot of it.

And whatever you do, call a performance discussion a check-in, not a review.

The oldest members of Gen Z are about to turn 30, yet companies are devoting more time and resources than ever to figuring out how to give this manager-befuddling generation better direction. For help, they are turning to a cottage industry of multigeneration-workplace consultants and even artificial-intelligence bots, while ripping up the script for what used to be once-a-year evaluations.

All of it is a departure for leaders who rose through the ranks in an era devoid of so much thought to effective coaching and criticism. “When I started my first job, I got a performance review a year later, and that was just expected,” said Adam Coyne, chief administrative officer at research and analytics firm Mathematica, which has shifted from annual reviews to quarterly, two-way check-ins for new junior hires.

“This is a group that wants a lot more real-time feedback,” added Coyne, 55.

It is a message managers say they are getting nonstop from surveys and all-hands meetings, not to mention the universities and colleges preparing graduates for the white-collar world of work: Used to the immediate validation of social-media likes and comments, even instantly posted grades, this generation of workers craves clear, frequent direction—and they feel disoriented and anxious when they don’t get it.

Gallup data suggest companies are still struggling to get the hang of it. Younger workers report some of the biggest drops in engagement at work over the past five years. Not knowing where they stand appears to be a big factor. The share of Gen Z and younger millennials who strongly agreed with the statement, “I know what is expected of me at work,” fell 9 points to 42% in surveys between 2020 and 2025.

That doesn’t mean they need the effusive praise that many managers claim they do, some 20-something workers say. “I personally dislike this style,” said Nathan Luckock, a 20-year-old engineer at an AI startup. More effective, he said, is just “pointing out mistakes and then offering a solution.”

That sounds familiar to Lindsey Pollak, a multigenerational workforce expert and executive coach, who says she coaches bosses to be as specific as possible. Instead of “be more responsive,” for instance, she suggests “need to hear from you within an hour of receiving an instruction.”

At Mathematica, Chief Executive Paul Decker said the firm switched to more frequent check-ins in part because so many new entry-level hires were peppering supervisors with questions like: “How am I doing?” and “What does the next level require?” At staff meetings, younger workers often questioned why things were done the way they had always been done.

That included things like “waiting months to learn whether you’re meeting expectations,” he said.

KPMG executives said they, too, began giving their younger workers more frequent assessments on skills like critical thinking and adaptability last year after interns said they wanted to hear more often how their skills were coming along. The firm wanted to “make sure that we scratch the itch,” said Jason LaRue, vice chair of talent and culture at KPMG’s U.S. practice.

Some managers are getting feedback on giving feedback from bots.

Joe Hirsch, a corporate speaker and author of “The Feedback Fix,” recently used an AI coaching platform to work with a tech-company manager on her delivery. She had been frustrated that one of her junior reports wasn’t grasping her pointers on pitching clients, so she role-played the conversation with the AI coach.

The problem, the bot advised, was that she wasn’t giving the employee enough context for why she wanted things done a certain way. “Let’s connect so I can share more about our approach and get your take on it,” it suggested she say.

That did the trick when she tried the approach in real life. “The advice finally landed,” Hirsch said.

Even a few, clear bullet points work, said Valerie Chapman, the 27-year-old founder and CEO of Ruth AI, a career strategist platform for women. “My generation likes to get feedback so that they know how they should adjust.”

Valerie Chapman headshot
Valerie Chapman, founder and CEO of Ruth AI. Dasha Semyonova

She recalls getting bullet-pointed direction when she worked as a strategic growth partner at real-estate company Compass a few years ago. The feedback started with praise before offering pointers.

“It would be, ‘Great job. Here are some things that you could do next week,’” she said. “If that comes in on a Friday, then my Gen Z brain knows exactly what I need to do on a Monday.”

Mike Ekbundit, director of GE Appliances’ engineering programs, said he has tried to make performance discussions with younger workers in rotational programs two-way dialogues rather than top-down critiques. So he revised online evaluations to include prompts for managers to ask questions like “Did you like the assignment leader?” and “Was it too much work?”

Managers are also asked to assess the program participants on nine different categories, like innovation and resilience, while employees are prompted to list their top strengths and weaknesses.

“I need high customer satisfaction to retain this highly sought-after talent, and this is part of how I get it,” he said.