{"id":663181,"date":"2025-04-17T00:33:48","date_gmt":"2025-04-16T21:33:48","guid":{"rendered":"https:\/\/en.buradabiliyorum.com\/openais-latest-ai-models-have-a-new-safeguard-to-prevent-biorisks\/"},"modified":"2025-04-17T00:33:48","modified_gmt":"2025-04-16T21:33:48","slug":"openais-latest-ai-models-have-a-new-safeguard-to-prevent-biorisks","status":"publish","type":"post","link":"https:\/\/buradabiliyorum.com\/en\/openais-latest-ai-models-have-a-new-safeguard-to-prevent-biorisks\/","title":{"rendered":"OpenAI&#8217;s latest AI models have a new safeguard to prevent biorisks"},"content":{"rendered":"<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">OpenAI says that it deployed a new system to monitor its latest AI reasoning models, o3 and o4-mini, for prompts related to biological and chemical threats. The system aims to prevent the models from offering advice that could instruct someone on carrying out potentially harmful attacks, <a rel=\"nofollow\" target=\"_blank\" rel=\"nofollow\" href=\"https:\/\/cdn.openai.com\/pdf\/2221c875-02dc-4789-800b-e7758f3722c1\/o3-and-o4-mini-system-card.pdf\">according to OpenAI\u2019s safety report<\/a>.<\/p>\n<p class=\"wp-block-paragraph\">O3 and o4-mini represent a meaningful capability increase over OpenAI\u2019s previous models, the company says, and thus pose new risks in the hands of bad actors. According to OpenAI\u2019s internal benchmarks, o3 is more skilled at answering questions around creating certain types of biological threats in particular. For this reason \u2014 and to mitigate other risks \u2014 OpenAI created the new monitoring system, which the company describes as a \u201csafety-focused reasoning monitor.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The monitor, custom-trained to reason about OpenAI\u2019s content policies, runs on top of o3 and o4-mini. It\u2019s designed to identify prompts related to biological and chemical risk and instruct the models to refuse to offer advice on those topics.<\/p>\n<p class=\"wp-block-paragraph\">To establish a baseline, OpenAI had red teamers spend around 1,000 hours flagging \u201cunsafe\u201d biorisk-related conversations from o3 and o4-mini. During a test in which OpenAI simulated the \u201cblocking logic\u201d of its safety monitor, the models declined to respond to risky prompts 98.7% of the time, according to OpenAI. <\/p>\n<p class=\"wp-block-paragraph\">OpenAI acknowledges that its test didn\u2019t account for people who might try new prompts after getting blocked by the monitor, which is why the company says it\u2019ll continue to rely in part on human monitoring.<\/p>\n<p class=\"wp-block-paragraph\">O3 and o4-mini don\u2019t cross OpenAI\u2019s \u201chigh risk\u201d threshold for biorisks, according to the company. However, compared to o1 and GPT-4, OpenAI says that early versions of o3 and o4-mini proved more helpful at answering questions around developing biological weapons.<\/p>\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1282\" height=\"588\" src=\"https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png?w=680\" alt=\"\" class=\"wp-image-2995283\" srcset=\"https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png 1282w, https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png?resize=150,69 150w, https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png?resize=300,138 300w, https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png?resize=768,352 768w, https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png?resize=680,312 680w, https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png?resize=1200,550 1200w, https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png?resize=1280,587 1280w, https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png?resize=430,197 430w, https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png?resize=720,330 720w, https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png?resize=900,413 900w, https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png?resize=800,367 800w, https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png?resize=668,306 668w, https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/04\/Screenshot-2025-04-16-at-2.09.10PM.png?resize=708,325 708w\" sizes=\"auto, (max-width: 1282px) 100vw, 1282px\"\/><figcaption class=\"wp-element-caption\"><span class=\"wp-element-caption__text\">Chart from o3 and o4-mini\u2019s system card (Screenshot: OpenAI)<\/span><\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">The company is actively tracking how its models could make it easier for malicious users to develop chemical and biological threats, according to OpenAI\u2019s recently updated <a rel=\"nofollow\" target=\"_blank\" rel=\"nofollow\" href=\"https:\/\/cdn.openai.com\/pdf\/18a02b5d-6b67-4cec-ab64-68cdfbddebcd\/preparedness-framework-v2.pdf\">Preparedness Framework<\/a>.<\/p>\n<p class=\"wp-block-paragraph\">OpenAI is increasingly relying on automated systems to mitigate the risks from its models. For example, to prevent <a rel=\"nofollow\" target=\"_blank\" rel=\"nofollow\" href=\"https:\/\/cdn.openai.com\/11998be9-5319-4302-bfbf-1167e093f1fb\/Native_Image_Generation_System_Card.pdf\">GPT-4o\u2019s native image generator from creating child sexual abuse material (CSAM)<\/a>, OpenAI says it uses on a reasoning monitor similar to the one the company deployed for o3 and o4-mini. <\/p>\n<p class=\"wp-block-paragraph\">Yet several researchers have raised concerns OpenAI isn\u2019t prioritizing safety as much as it should. One of the company\u2019s red-teaming partners, Metr, said it had relatively little time to test o3 on a benchmark for deceptive behavior. Meanwhile, OpenAI decided not to release a safety report for its GPT-4.1 model, which launched earlier this week.<\/p>\n<\/div>\n<blockquote><p><strong><span style=\"color: #ff6600;\">If you liked the article, do not forget to share it with your friends. Follow us on\u00a0<span style=\"color: #ff0000;\"><a style=\"color: #ff0000;\" href=\"https:\/\/news.google.com\/publications\/CAAqBwgKMN63nwsw68G3Aw\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">Google News<\/a><\/span>\u00a0too, click on the star and choose us from your favorites.<\/span><\/strong><\/p><\/blockquote>\n<blockquote>\n<p style=\"text-align: center;\"><strong>If you want to read more like this article, you can visit our <span style=\"color: #ff9900;\"><a style=\"color: #ff9900;\" href=\"https:\/\/en.buradabiliyorum.com\/category\/technology\/\" target=\"_blank\" >Technology<\/a><\/span> category.<\/strong><\/p>\n<\/blockquote>\n<p><span style=\"color: black;\"><a style=\"color: #ff9900;\" href=\"https:\/\/techcrunch.com\/2025\/04\/16\/openais-latest-ai-models-have-a-new-safeguard-to-prevent-biorisks\/\" target=\"_blank\" >Source<\/a><\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>OpenAI says that it deployed a new system to monitor its latest AI reasoning models, o3 and o4-mini, for prompts related to biological and chemical threats. The system aims to prevent the models from offering advice that could instruct someone on carrying out potentially harmful attacks, according to OpenAI\u2019s safety report. O3 and o4-mini represent&#8230;<\/p>\n","protected":false},"author":1,"featured_media":663182,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/01\/GettyImages-2191707579.jpg?w=1024","fifu_image_alt":"","footnotes":""},"categories":[18],"tags":[77337,152633,138467],"class_list":["post-663181","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-technology","tag-ai","tag-ai-safety","tag-chatgpt"],"_links":{"self":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/posts\/663181","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/comments?post=663181"}],"version-history":[{"count":0,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/posts\/663181\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/media\/663182"}],"wp:attachment":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/media?parent=663181"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/categories?post=663181"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/tags?post=663181"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}