{"id":741789,"date":"2026-07-30T10:15:23","date_gmt":"2026-07-30T07:15:23","guid":{"rendered":"https:\/\/buradabiliyorum.com\/en\/aligning-the-model-was-never-going-to-govern-it\/"},"modified":"2026-07-30T10:15:23","modified_gmt":"2026-07-30T07:15:23","slug":"aligning-the-model-was-never-going-to-govern-it","status":"publish","type":"post","link":"https:\/\/buradabiliyorum.com\/en\/aligning-the-model-was-never-going-to-govern-it\/","title":{"rendered":"Aligning the model was never going to govern it"},"content":{"rendered":"<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_87 counter-hierarchy ez-toc-counter ez-toc-custom ez-toc-container-direction\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<label for=\"ez-toc-cssicon-toggle-item-6a9fc3e07962d\" class=\"ez-toc-cssicon-toggle-label\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #dd3333;color:#dd3333\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #dd3333;color:#dd3333\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/label><input type=\"checkbox\"  id=\"ez-toc-cssicon-toggle-item-6a9fc3e07962d\" checked aria-label=\"Toggle\" \/><nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/buradabiliyorum.com\/en\/aligning-the-model-was-never-going-to-govern-it\/#Safety_is_not_the_same_as_governance\" >Safety is not the same as governance<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/buradabiliyorum.com\/en\/aligning-the-model-was-never-going-to-govern-it\/#The_problem_is_bigger_than_the_model_and_it_is_growing\" >The problem is bigger than the model, and it is growing<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/buradabiliyorum.com\/en\/aligning-the-model-was-never-going-to-govern-it\/#Why_aligning_the_model_cannot_fix_it\" >Why aligning the model cannot fix it<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/buradabiliyorum.com\/en\/aligning-the-model-was-never-going-to-govern-it\/#The_fix_is_a_layer_around_the_model\" >The fix is a layer around the model<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/buradabiliyorum.com\/en\/aligning-the-model-was-never-going-to-govern-it\/#The_right_layer\" >The right layer<\/a><\/li><\/ul><\/nav><\/div>\n<div id=\"article-main-content\">\n<p>Most AI roadmaps rest on a quiet bet: that the labs will eventually ship a model safe and aligned enough to simply trust in production. Better training, better guardrails, one more version, and the thing behaves.<\/p>\n<p>Here is the flaw in that bet. Even a perfectly aligned model cannot tell you who used it, on what data, under whose policy, or hand you a record you could show an auditor. Those are not facts about how the model behaves.<\/p>\n<p>They are facts about how it was deployed, and they live entirely outside the weights. A better model answers a different question than the one regulators, auditors, and security teams are actually asking, and the gap between those two questions is where the real risk now sits.<\/p>\n<p>Security engineering named this trap fifty years ago. A 1972 US Air Force study defined the <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/csrc.nist.gov\/files\/pubs\/conference\/1998\/10\/08\/proceedings-of-the-21st-nissc-1998\/final\/docs\/early-cs-papers\/ande72a.pdf\" target=\"_blank\" rel=\"nofollow noopener\">reference monitor<\/a>, the component that decides whether an action is allowed, and set three conditions for trusting it: it must be tamperproof, invoked on every access, and small enough to be completely verified.<\/p>\n<div class=\"inarticle-wrapper latest channel-cta hs-embed-tnw\">\n<div id=\"hs-embed-tnw\" class=\"channel-cta-wrapper\">\n<div class=\"channel-cta-img\"><img decoding=\"async\" class=\"js-lazy\" src=\"https:\/\/media.thenextweb.com\/hardfork-2018\/uploads\/visuals\/tnw-newsletter.png\"\/><\/div>\n<p><img decoding=\"async\" src=\"https:\/\/media.thenextweb.com\/hardfork-2018\/uploads\/visuals\/tnw-newsletter.png\"\/><\/p>\n<div class=\"channel-cta-input\">\n<p class=\"channel-cta-title\">The \ud83d\udc9c of EU tech<\/p>\n<p class=\"channel-cta-tagline\">The latest rumblings from the EU tech scene, a story from our wise ol&#8217; founder Boris, and some questionable AI art. It&#8217;s free, every week, in your inbox. Sign up now!<\/p>\n<\/div>\n<\/div>\n<\/div>\n<p>A modern frontier model is none of the three. Alignment tries to make the model enforce the rules meant to constrain it, from the inside, and a system cannot be the thing that governs itself. Governance has to sit around the model, not in it.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Safety_is_not_the_same_as_governance\"><\/span>Safety is not the same as governance<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p style=\"font-weight: 400;\">The confusion hides in one word. Safety and alignment ask whether a model tends to behave well. That is a disposition, it lives in the weights, and the labs have gotten genuinely good at shaping it. Governance asks something different: who used which model, on what data, under whose policy, and with what auditable record.<\/p>\n<p style=\"font-weight: 400;\">That is a property of a specific deployment. Training can shape the first. It cannot, by construction, supply the second. Treating them as one problem is the category error underneath most AI risk conversations.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"The_problem_is_bigger_than_the_model_and_it_is_growing\"><\/span>The problem is bigger than the model, and it is growing<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p style=\"font-weight: 400;\">This is not a niche concern for the labs. It lands on everyone who deploys AI, and it is expanding for two reasons.<\/p>\n<p style=\"font-weight: 400;\">First, models are gaining agency. When a model only produced text, a bad interaction meant a bad answer. Now models take actions through tools and protocols like <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/modelcontextprotocol.io\/specification\/2025-06-18\/basic\/authorization\" target=\"_blank\" rel=\"nofollow noopener\">MCP<\/a>, and every tool call is a fresh decision about who is acting, on what data, under whose authority. Those decisions multiply with every task and every user.<\/p>\n<p style=\"font-weight: 400;\">OWASP\u2019s 2026 tracking finds that most of the agentic projects it follows are coding agents, and that prompt injection maps to six of the ten risks on its agentic top-ten list. The ungoverned surface grows as fast as adoption.<\/p>\n<p style=\"font-weight: 400;\">Second, the accountability is now legal. California\u2019s SB 53 and the European Union\u2019s rules for <a href=\"https:\/\/buradabiliyorum.com\/en\/category\/general\/\" data-internallinksmanager029f6b8e52c=\"3\" title=\"General\" target=\"_blank\" rel=\"noopener\">general<\/a>-purpose AI put obligations on the labs, but the EU AI Act also places separate <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/artificialintelligenceact.eu\/article\/26\/\" target=\"_blank\" rel=\"nofollow noopener\">duties on deployers<\/a>, and data-protection law already holds organizations responsible for how personal data is used.<\/p>\n<p style=\"font-weight: 400;\">When a regulator asks what your AI did, \u201cwe used an aligned model\u201d is not an answer, and a vendor\u2019s safety report will not tell them who inside your company sent which data to which model last Tuesday.<\/p>\n<p style=\"font-weight: 400;\">Public benchmarks do not help either: StrongREJECT, HarmBench, and AgentHarm grade the model on generic prompts, not your users, your data, or your policy on a given day. A high score is a property of the model, not an audit trail.<\/p>\n<p style=\"font-weight: 400;\">So the scope is wide and concrete: every organization that deploys AI, a surface that grows with every new agent, and a liability that is now enforceable. None of it is a model-quality problem.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Why_aligning_the_model_cannot_fix_it\"><\/span>Why aligning the model cannot fix it<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p style=\"font-weight: 400;\">Go back to the 1972 test, because it shows exactly where the fix cannot live.<\/p>\n<p style=\"font-weight: 400;\">A model is not tamperproof. <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/simonwillison.net\/2022\/Sep\/12\/prompt-injection\/\" target=\"_blank\" rel=\"nofollow noopener\">Prompt injection<\/a>, named in 2022, works because instructions and data ride the same channel of text, so a document the model reads can overwrite the instruction it was given.<\/p>\n<p style=\"font-weight: 400;\">Four years on it is unsolved: OWASP <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/genai.owasp.org\/llmrisk\/llm01-prompt-injection\/\" target=\"_blank\" rel=\"nofollow noopener\">says<\/a> it is unclear whether fool-proof prevention exists, and a 2026 paper <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/arxiv.org\/abs\/2605.17634\" target=\"_blank\" rel=\"nofollow noopener\">argues<\/a> agents may always be vulnerable. A control you can rewrite with the input it is inspecting is not a control.<\/p>\n<p style=\"font-weight: 400;\">A model is not reliably invoked as a gate. Give it tools and it becomes a textbook confused deputy, a 1988 term for a trusted program that an untrusted caller tricks into misusing its authority, because it cannot tell an instruction from the data it is reading.<\/p>\n<p style=\"font-weight: 400;\">A model is not verifiable. You cannot inspect billions of opaque parameters the way you can audit a small enforcement kernel.<\/p>\n<p style=\"font-weight: 400;\">Fail all three tests and the conclusion is not \u201calign harder.\u201d It is that the model cannot be the enforcement mechanism, so the enforcement has to live somewhere else.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"The_fix_is_a_layer_around_the_model\"><\/span>The fix is a layer around the model<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p style=\"font-weight: 400;\">If the model cannot govern itself, the answer is not a better model but a thin control layer around it, one that treats the model as an untrusted component and <a href=\"https:\/\/buradabiliyorum.com\/en\/category\/social-mediaa\/\" data-internallinksmanager029f6b8e52c=\"1\" title=\"Social Media\" target=\"_blank\" rel=\"noopener\">media<\/a>tes everything it tries to do.<\/p>\n<p style=\"font-weight: 400;\">This is not new architecture to invent. It is the 1972 reference monitor rebuilt for an agent, and here is how it works.<\/p>\n<p style=\"font-weight: 400;\">Picture a request moving through it. The model reads its input and proposes an action: call this API, write this file, send this message.<\/p>\n<p style=\"font-weight: 400;\">That proposal is never trusted on its own. It leaves the model as a request, not a command, and passes to a small policy engine outside the model, sitting in the path between the agent and the resource it wants to reach, the way <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/nvlpubs.nist.gov\/nistpubs\/specialpublications\/NIST.SP.800-207.pdf\" target=\"_blank\" rel=\"nofollow noopener\">zero-trust architecture<\/a> separates the point that decides whether an action is allowed from the point that carries it out.<\/p>\n<p style=\"font-weight: 400;\">The engine checks the action against two things the weights never had: your actual policy, and the narrow set of capabilities this agent was granted for this task.<\/p>\n<p style=\"font-weight: 400;\">If the action clears both, an enforcement point performs it; if not, it is refused. Either way, who acted, through which model, on what data, under which policy, and what was decided is written to a log.<\/p>\n<p style=\"font-weight: 400;\">That one flow restores the three properties the model could not provide. It is tamperproof, because the policy lives in a small external engine rather than in text the model can be argued out of. It is always invoked, because nothing reaches a real resource without passing through it.<\/p>\n<p style=\"font-weight: 400;\">And it is verifiable, because an enforcement point small enough to read is one you can actually trust, in a way billions of opaque weights never can be. This is why the strongest research defenses already work this way: CaMeL, from Google DeepMind and ETH Zurich, gets provable security by containing the model exactly like this, and it holds even when the underlying model stays vulnerable. The guarantee lives in the architecture, not the weights.<\/p>\n<p style=\"font-weight: 400;\">None of this makes the model safe, and that is the point. It makes the surrounding system accountable, which is a smaller and far more achievable goal, and it is the one the deployment problem actually demands.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"The_right_layer\"><\/span>The right layer<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p style=\"font-weight: 400;\">This is not an argument against alignment. Alignment is necessary, and it is improving. It is simply the wrong layer for the question that regulators, auditors, and users are actually asking. The answer to that question has been on the shelf since 1972, and the industry keeps re-deriving it the expensive way.<\/p>\n<p style=\"font-weight: 400;\">The next model will be safer than this one. It still will not know who used it, on what data, or under whose policy. Closing that gap is the work, and it is the story.<\/p>\n<\/p><\/div>\n<blockquote><p><strong><span style=\"color: #ff6600;\">If you liked the article, do not forget to share it with your friends. Follow us on\u00a0<span style=\"color: #ff0000;\"><a style=\"color: #ff0000;\" href=\"https:\/\/news.google.com\/publications\/CAAqBwgKMN63nwsw68G3Aw\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">Google News<\/a><\/span>\u00a0too, click on the star and choose us from your favorites.<\/span><\/strong><\/p><\/blockquote>\n<blockquote>\n<p style=\"text-align: center;\"><strong>If you want to read more like this article, you can visit our <span style=\"color: #ff9900;\"><a style=\"color: #ff9900;\" href=\"https:\/\/buradabiliyorum.com\/en\/category\/technology\/\" target=\"_blank\" >Technology category.<\/a><\/span><\/strong><\/p>\n<\/blockquote>\n<p><span style=\"color: black;\"><a style=\"color: #ff9900;\" href=\"https:\/\/thenextweb.com\/news\/aligning-the-model-was-never-going-to-govern-it\" target=\"_blank\" >Source<\/a><\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Most AI roadmaps rest on a quiet bet: that the labs will eventually ship a model safe and aligned enough to simply trust in production. Better training, better guardrails, one more version, and the thing behaves. Here is the flaw in that bet. Even a perfectly aligned model cannot tell you who used it, on&#8230;<\/p>\n","protected":false},"author":1,"featured_media":741790,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/media.thenextweb.com\/2026\/07\/AI-models.avif","fifu_image_alt":"","footnotes":""},"categories":[18],"tags":[],"class_list":["post-741789","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-technology"],"_links":{"self":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/posts\/741789","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/comments?post=741789"}],"version-history":[{"count":0,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/posts\/741789\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/media\/741790"}],"wp:attachment":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/media?parent=741789"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/categories?post=741789"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/tags?post=741789"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}