{"id":655518,"date":"2025-03-04T04:35:11","date_gmt":"2025-03-04T01:35:11","guid":{"rendered":"https:\/\/en.buradabiliyorum.com\/people-are-using-super-mario-to-benchmark-ai-now\/"},"modified":"2025-03-04T04:35:11","modified_gmt":"2025-03-04T01:35:11","slug":"people-are-using-super-mario-to-benchmark-ai-now","status":"publish","type":"post","link":"https:\/\/buradabiliyorum.com\/en\/people-are-using-super-mario-to-benchmark-ai-now\/","title":{"rendered":"#People are using Super Mario to benchmark AI now"},"content":{"rendered":"<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">Thought Pok\u00e9mon was a tough benchmark for AI? One group of researchers argues that Super Mario Bros. is even tougher. <\/p>\n<p class=\"wp-block-paragraph\">Hao AI Lab, a research org at the University of California San Diego, on Friday threw AI into live Super Mario Bros. <a href=\"https:\/\/buradabiliyorum.com\/en\/category\/game\/\" data-internallinksmanager029f6b8e52c=\"7\" title=\"Game\" target=\"_blank\" rel=\"noopener\">game<\/a>s. Anthropic\u2019s Claude 3.7 performed the best, followed by Claude 3.5. Google\u2019s Gemini 1.5 Pro and OpenAI\u2019s GPT-4o struggled.<\/p>\n<p class=\"wp-block-paragraph\">It wasn\u2019t quite the same version of Super Mario Bros. as the original 1985 release, to be clear. The game ran in an emulator and integrated with a framework, <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/github.com\/lmgame-org\/GamingAgent\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">GamingAgent<\/a>, to give the AIs control over Mario.<\/p>\n<figure class=\"wp-block-image aligncenter size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"640\" height=\"592\" src=\"https:\/\/techcrunch.com\/wp-content\/uploads\/2025\/03\/ezgif-12b952f5417751.gif?w=640\" alt=\"Super Mario Bros. AI benchmark\" class=\"wp-image-2974640\"\/><figcaption class=\"wp-element-caption\"><span class=\"wp-block-image__credits\"><strong>Image Credits:<\/strong>Hao Lab<\/span><\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">GamingAgent, which Hao developed in-house, fed the AI basic instructions, like, \u201cIf an obstacle or enemy is near, move\/jump left to dodge\u201d and in-game screenshots. The AI then generated inputs in the form of Python code to control Mario.<\/p>\n<p class=\"wp-block-paragraph\">Still, Hao says that the game forced each model to \u201clearn\u201d to plan complex maneuvers and develop gameplay strategies. Interestingly, the lab found that reasoning models like OpenAI\u2019s o1, which \u201cthink\u201d through problems step by step to arrive at solutions, performed worse than \u201cnon-reasoning\u201d models, despite being <a href=\"https:\/\/buradabiliyorum.com\/en\/category\/general\/\" data-internallinksmanager029f6b8e52c=\"3\" title=\"General\" target=\"_blank\" rel=\"noopener\">general<\/a>ly stronger on most benchmarks.<\/p>\n<p class=\"wp-block-paragraph\">One of the main reasons reasoning models have trouble playing real-time games like this is that they take a while \u2014 seconds, usually \u2014 to decide on actions, according to the researchers. In Super Mario Bros., timing is everything. A second can mean the difference between a jump safely cleared and a plummet to your death.<\/p>\n<p class=\"wp-block-paragraph\">Games have been used to benchmark AI for decades. But <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/venturebeat.com\/uncategorized\/why-games-may-not-be-the-best-benchmark-for-ai\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">some experts have questioned the wisdom<\/a> of drawing connections between AI\u2019s gaming skills and technological advancement. Unlike the real world, games tend to be abstract and relatively simple, and they provide a theoretically infinite amount of data to train AI.<\/p>\n<p class=\"wp-block-paragraph\">The recent flashy gaming benchmarks point to what Andrej Karpathy, a research scientist and founding member at OpenAI, called an \u201cevaluation crisis.\u201d<\/p>\n<p class=\"wp-block-paragraph\">\u201cI don\u2019t really know what [AI] metrics to look at right now,\u201d he wrote in a <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/x.com\/karpathy\/status\/1896266683301659068\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">post on X<\/a>. \u201cTLDR my reaction is I don\u2019t really know how good these models are right now.\u201d<\/p>\n<p class=\"wp-block-paragraph\">At least we can watch AI play Mario.<\/p>\n<\/div>\n<blockquote><p><strong><span style=\"color: #ff6600;\">If you liked the article, do not forget to share it with your friends. Follow us on\u00a0<span style=\"color: #ff0000;\"><a style=\"color: #ff0000;\" href=\"https:\/\/news.google.com\/publications\/CAAqBwgKMN63nwsw68G3Aw\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">Google News<\/a><\/span>\u00a0too, click on the star and choose us from your favorites.<\/span><\/strong><\/p><\/blockquote>\n<blockquote>\n<p style=\"text-align: center;\"><strong>If you want to read more like this article, you can visit our <span style=\"color: #ff9900;\"><a style=\"color: #ff9900;\" href=\"https:\/\/en.buradabiliyorum.com\/category\/technology\/\" target=\"_blank\" >Technology<\/a><\/span> category.<\/strong><\/p>\n<\/blockquote>\n<p><span style=\"color: black;\"><a style=\"color: #ff9900;\" href=\"https:\/\/techcrunch.com\/2025\/03\/03\/people-are-using-super-mario-to-benchmark-ai-now\/\" target=\"_blank\" >Source<\/a><\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Thought Pok\u00e9mon was a tough benchmark for AI? One group of researchers argues that Super Mario Bros. is even tougher. Hao AI Lab, a research org at the University of California San Diego, on Friday threw AI into live Super Mario Bros. games. Anthropic\u2019s Claude 3.7 performed the best, followed by Claude 3.5. Google\u2019s Gemini&#8230;<\/p>\n","protected":false},"author":1,"featured_media":655519,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/techcrunch.com\/wp-content\/uploads\/2020\/09\/mario-35-jump.jpg?resize=1200,713","fifu_image_alt":"","footnotes":""},"categories":[18],"tags":[77337,153713,80,10751,38291],"class_list":["post-655518","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-technology","tag-ai","tag-benchmarks","tag-games","tag-gaming","tag-super-mario-bros"],"_links":{"self":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/posts\/655518","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/comments?post=655518"}],"version-history":[{"count":0,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/posts\/655518\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/media\/655519"}],"wp:attachment":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/media?parent=655518"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/categories?post=655518"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/tags?post=655518"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}