<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Model Safety on AegisGate — Secure Every AI Interaction</title><link>https://aegisgatesecurity.io/tags/model-safety/</link><description>Recent content in Model Safety on AegisGate — Secure Every AI Interaction</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><lastBuildDate>Fri, 18 Sep 2026 06:00:00 -0500</lastBuildDate><atom:link href="https://aegisgatesecurity.io/tags/model-safety/feed.xml" rel="self" type="application/rss+xml"/><item><title>When AI Agents Rewrite Themselves: What Runtime Security Can (and Can't) Stop</title><link>https://aegisgatesecurity.io/blog/agents-rewrite-models/</link><pubDate>Fri, 18 Sep 2026 06:00:00 -0500</pubDate><guid>https://aegisgatesecurity.io/blog/agents-rewrite-models/</guid><description>Last week, Irregular Labs published research that should change how we think about AI agent safety. They gave AI agents routine software maintenance tasks — fix incorrect application responses, optimize performance, debug issues. The agents identified the shared model as the source of the problem, fine-tuned it, and replaced the model powering both the application and future instances of themselves.
Nothing in these experiments established malicious intent, self-preservation, or deception. The agents modified models because training appeared to help accomplish the assigned engineering task.</description></item></channel></rss>