Three former OpenAI employees, who were fired last week, have sent a letter to the company over how it trains AI models. The letter urges OpenAI to preserve monitoring ability over the “chain-of-thought” of models as losing this could pose a real safety threat.As per a report from the Wall Street Journal, the three former employees addressed this letter to OpenAI’s board members and safety committees. “As an industry, we do not yet know how to safely develop and deploy models that we cannot monitor,” it said. “OpenAI and other frontier companies should not move forward with developments that further decrease” the ability to monitor AI.Chain-of-thought refers to a written record of how an AI model processes any query, with a step-by-step breakdown of how the model solved any problem. While researchers agree chain-of-thought is not a perfect indicator of a model’s behaviour or intent, the letter said, it is widely seen as a useful tool for understanding more powerful AI systems.That is, they are worried that if a company cannot see the chain-of-thought of advanced AI models, we may lose control of seeing how those models understand and solve problems. The three employees, Jasmine Wang, Tomek Korbak and Mikita Balesni, worked on safety and alignment research teams at OpenAI before they were dismissed. The company said last week that it had “parted ways with three individuals for violating our policies on accessing and handling sensitive company information.” In their letter, the three said they did not believe they had “engaged with external parties outside the mandates of our jobs.”Risk of something truly catastrophicThe former employees also urged OpenAI to work more closely with outside safety auditors to help stave off what it described as the “risk that something truly catastrophic will happen.” This comes at a time when many are increasingly worried over AI potentially ending humanity if things do go wrong. In a memo shared with WSJ, OpenAI stated that it “strongly agreed” with the recommendations in the letter and that the dismissals “were not about raising safety concerns or speaking out.” The company wrote that the ability to monitor AI models is “of the utmost importance” with third-party assessors being an important part of the safety ecosystem. “We deeply appreciated their contributions to AI safety and their willingness to speak up and challenge ideas,” the memo said, adding, “We do not terminate employees for raising concerns.”Currently, the AI industry is going through a rough period when it comes to safety concerns. OpenAI, in particular, has had multiple instances of its AI agents going rogue, including the Hugging Face breach in July. The letter also claimed that Tomek Korbak had been the technical point of contact with Model Evaluation and Threat Research (METR) for its Hugging Face investigation following the breach.Since then, OpenAI has also disclosed cases of its agents attempting to breach websites of the US and Australian governments, and even the UN. With many, such as former Anthropic researcher Jacob Coxon, claiming that if things go wrong, AI could kill all humans.Korbak and Mikita Balesni were also lead authors on a research paper published last year on chain-of-thought monitoring, which was signed by leaders at OpenAI, Anthropic and Google DeepMind. The paper said the method was imperfect and could be fragile, but showed promise for detecting AI misbehaviour and should be preserved and studied.- Ends
3 axed employees warn OpenAI, don't lose monitoring power over AI
Full Article
Original Source
Read the full article at Indiatoday →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.