---
title: What Keeping an AI Agent Running Requires – AI Abstraction
canonical_url: https://ai-abstraction.com/ai-agent-maintenance-security-and-compliance/
description: Model retirements, prompt injection, the FTC Safeguards Rule and NYDFS Part 500: what it takes to run an AI agent safely in a regulated practice.
last_modified: 2026-08-21
---

# What Keeping an AI Agent Running Requires

[Security Compliance](/category/security-compliance/)

By [Jay Taheri](/about/#member-jay-taheri)

August 21, 2026

13 min read

![An office manager working through a printed checklist beside a laptop](/assets/media/ai-agent-maintenance-security-and-compliance-image.jpg)

Image generated with AI.

Most conversations about AI agents stop at launch. The agent answers the phone, drafts the follow-up, updates the record, and everyone agrees it works. That is a real milestone and worth celebrating.

It is also the point at which the interesting problems start.

An agent is not a website that sits still once you publish it. It sits on top of models that get retired on a schedule, connects to systems that change without asking you, and, if you are a mortgage broker or a CPA firm, handles exactly the category of customer information that federal and state regulators have specific written rules about. None of that is a reason to avoid building one. It is a reason to know what you are signing up for before you do.

Here is what the four ongoing jobs actually involve, and an honest account of which parts a managed service takes off your desk and which parts it cannot.

## Maintenance: the parts underneath your agent have expiry dates

The most common misconception about AI agents is that maintenance means occasionally improving them. It does include that. But the larger share is keeping the thing working at all while the ground moves underneath it.

### Model retirement is scheduled, not hypothetical

The AI model your agent calls is a product with a published end of life, and the dates are not far apart. Anthropic, to take one provider that [publishes its schedule openly](https://platform.claude.com/docs/en/about-claude/model-deprecations), gives at least 60 days’ notice before retiring a model and distinguishes between deprecated, meaning still running but no longer recommended, and retired, meaning “requests to retired models will fail.”

Look at what that meant in practice. Six Claude models reached their retirement date inside the last twelve months: two in February 2026, one in April, two in June, and one in August. Any agent still pointed at those model names stopped working on those dates, not gradually, but at once. Every other major provider runs a comparable schedule.

Migrating is rarely a matter of changing a string. A newer model is usually better, but it is also differently behaved. Prompts tuned against the old model produce slightly different output against the new one. That difference can be an improvement, or it can be your intake agent quietly starting to phrase a disclosure in a way your compliance officer would not have approved. You only find out by testing before you switch, against real examples, not three obvious ones.

### The integrations move too

Your CRM, your calendar, your phone system, and your document storage all publish their own APIs, and they change them on their own timetables. An authentication method gets retired. A field gets renamed. A rate limit gets tightened. Each of those is small on its own. Collectively they are the reason an agent that worked perfectly in March is dropping every third lead in September.

### Silent failure is the normal failure mode

This is the part that matters most and gets discussed least. When a website goes down, you know. When an agent degrades, it usually keeps responding. It just responds worse. It stops picking up an attachment, or starts summarizing three fields instead of five, or answers a question it should have escalated.

Nobody files a ticket about an agent that is 90 percent right, because the 10 percent is invisible from the inside. It becomes visible when a client notices. That gap, between the day something breaks and the day someone tells you, is the single strongest practical argument for continuous monitoring rather than periodic checking.

## Security: an agent is a genuinely new attack surface

The security industry has spent two years cataloguing how these systems fail, and the resulting list is worth knowing about even if you never read a security document again. The [OWASP Top 10 for LLM Applications](https://genai.owasp.org/llm-top-10/) is the reference most practitioners work from. Four of its ten entries matter directly to a small business running an agent on real customer data.

### Prompt injection, which is not a solved problem

Prompt injection sits at number one on that list, and it deserves the position. The short version: your agent reads text, and text can contain instructions. If your agent processes an inbound email, and that email contains a line telling the agent to forward the thread elsewhere or ignore its previous instructions, a naively built agent may simply do it.

This is not a bug that gets patched once. It is a structural property of systems that take instructions in the same channel they take data. It is managed with layered defenses, not eliminated. Anyone who tells you their agent is immune to prompt injection is either not being straight with you or has not looked closely.

### Excessive agency

Number six on the same list is excessive agency: giving the agent more permission than the job requires. An agent that only needs to read your calendar should not hold write access to your CRM. An agent that drafts client emails should not also hold the ability to send them.

This is the most preventable risk on the list and the most commonly ignored, because broad permissions make the build easier. Narrow permissions take longer to configure and mean the worst-case outcome of a successful injection is an agent that does something useless rather than something expensive.

### Sensitive information disclosure and misinformation

The other two worth knowing: an agent can surface data to someone who should not see it, and it can state something false with complete confidence. For a CPA firm, the second one has a specific shape. An agent that invents a filing deadline sounds exactly as authoritative as one that reports the correct deadline.

The defenses are unglamorous. Ground the agent in your actual records rather than the model’s memory. Require it to say when it does not know. Keep anything client-facing behind a person’s approval. None of that is clever. All of it works.

## Compliance: you are probably already covered by a rule you have not read

This is where the conversation gets specific to your industry, and where a surprising number of firms discover they have been subject to a written requirement for years.

### Mortgage brokers and tax preparers are “financial institutions” under federal law

The [FTC Safeguards Rule](https://www.ftc.gov/business-guidance/resources/ftc-safeguards-rule-what-your-business-needs-know) applies to non-bank financial institutions, and the FTC’s own list of who that covers includes **mortgage lenders and brokers** and **tax preparation firms** by name. If you are either, you are required to have a written information security program with nine specific elements: a designated Qualified Individual, a documented risk assessment, implemented safeguards including access controls and encryption and multi-factor authentication, monitoring and testing, staff training, service provider oversight, periodic updates, a written incident response plan, and an annual report from the Qualified Individual to your board or governing body.

An AI agent that reads client email and touches loan or tax records lands squarely inside the scope of that program. It is a system holding customer information, so it belongs in the risk assessment, the access controls, and the incident response plan.

### New York licensees have a second, stricter regime

If you hold a New York license, [23 NYCRR Part 500](https://www.dfs.ny.gov/cybersecurity/23-NYCRR-Part-500) applies on top. A covered entity there is “any person operating under or required to operate under a license, registration, charter, certificate, permit, accreditation or similar authorization under the Banking Law, the Insurance Law or the Financial Services Law.” That sweeps in a great many Bergen County firms that do business across the river.

Part 500 is more prescriptive than the federal rule in two places that matter here. Section 500.11 requires written policies covering “the identification and risk assessment of third-party service providers,” the “minimum cybersecurity practices required to be met by such third-party service providers,” due diligence before you engage them, and “periodic assessment of such third-party service providers based on the risk they present.” Section 500.17 requires notifying the superintendent of a cybersecurity incident “as promptly as possible but in no event later than 72 hours” after you determine one occurred.

Seventy-two hours is not very long if you are trying to reconstruct what an agent did from memory.

There is a small-firm exemption, and it is worth reading carefully rather than assuming it covers you. Section 500.19(a) grants a limited exemption to a covered entity with fewer than 20 employees and contractors, or under $7.5 million in gross annual revenue across the last three fiscal years, or under $15 million in year-end assets. That describes a lot of Bergen County practices.

But look at what the limited exemption actually exempts you from: sections 500.4, 500.5, 500.6, 500.8, 500.10, parts of 500.14, 500.15 and 500.16. Section 500.11 is not on that list. Neither is section 500.17.

In other words, the two requirements that bear most directly on hiring an AI vendor, the written third-party oversight policy and the 72-hour incident notice, are the two a small firm does not get relief from. If you are a ten-person shop with a New York license who assumed the exemption made this someone else’s problem, it is worth checking that assumption against the text.

### Hiring a vendor does not transfer the obligation

This is the point most worth internalizing, and it cuts against the vendor’s own interest to say plainly.

Bringing in an outside provider does not move your regulatory duty onto them. The FTC’s guidance on service providers is direct about the shape of the responsibility that stays with you: “Your contracts must spell out your security expectations, build in ways to monitor your service provider’s work, and provide for periodic reassessments of their suitability.” The obligation to oversee is yours. You cannot outsource it by outsourcing the work.

New Jersey has said something structurally similar about discrimination. The [Attorney General’s January 2025 guidance on algorithmic discrimination](https://www.njoag.gov/attorney-general-platkin-and-division-on-civil-rights-announce-new-guidance-on-algorithmic-discrimination/) states that “a covered entity is not shielded from liability for algorithmic discrimination that results from the entity’s use of an automated decision-making tool simply because the tool was developed by a third party,” and adds that an entity “can violate the LAD even if it has no intent to discriminate.” If an agent is deciding which inquiries get a callback and which do not, that guidance is about you, and good intentions are not a defense.

So the honest framing is this: a managed provider does not make you compliant. A managed provider gives you the artifacts your compliance program needs, and the accountable party is still you.

## Regulation: how to stay ready without guessing

Nobody can tell you with confidence what AI regulation looks like in three years. What is available now is a widely recognized structure for organizing the question.

[NIST’s AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework) is voluntary, not binding, and organized around four functions: govern, map, measure, and manage. In July 2024 NIST published a companion Generative AI Profile, NIST AI 600-1, addressing the risks specific to systems like these. It is worth knowing about not because anyone will fine you for ignoring it, but because it is the vocabulary regulators, insurers, and enterprise clients are converging on. A firm that can already say who governs its agent, what data it touches, how its behavior is measured, and what happens when it misbehaves is well positioned for whatever specific rule arrives.

The practical advice is boring and durable: keep records. The single artifact that satisfies the largest number of current and plausible future requirements is a complete, timestamped log of what the agent did, what it drew on, who approved it, and when. Firms that keep that log adapt to new rules by writing a report. Firms that do not adapt by starting over.

## What “managed” actually changes, and what it does not

Being specific here matters more than being enthusiastic.

### What it should take off your desk

A managed service means the model migration is somebody’s scheduled job rather than your emergency. It means integration breakage is caught by monitoring rather than by a client. It means prompt changes get tested before they ship, permissions get reviewed rather than accumulating, and the activity log exists by default rather than being reconstructed after an incident. In our own service, that is uptime monitoring, model upgrades, monthly prompt tuning, security patching, integration maintenance, and a named person who answers when something breaks.

It also means that when your Qualified Individual writes the annual report, or a New York examiner asks about third-party oversight, there is something to hand over.

### What it does not do

It does not make you compliant. It does not make you the party who stops being responsible. It does not eliminate prompt injection, and it does not mean nothing will ever go wrong.

What it changes is who is watching, how fast the problem is found, and whether there is a record of it afterward. For most small firms, that is the entire difference between an agent that is an asset and one that is a liability nobody is looking at.

## Five questions to ask any provider, including us

Skip the pitch and ask these:

1.  **What happens on the day the model we are using is retired?** A good answer includes a testing process and who pays for the migration. A bad answer treats it as hypothetical.
2.  **What can the agent do without a person approving it?** Ask for the actual permission list, not a reassurance.
3.  **Show me the record of what the system did last Tuesday.** Not a dashboard. The log: which record, what changed, who approved it, when.
4.  **What do you give me for my written information security program?** If they do not know what that phrase means and you are a broker or a tax preparer, keep looking.
5.  **If we leave, what comes with us?** Your data and your activity records should be yours without argument. Whether prompts and configuration travel varies by provider, so ask rather than assume.

## The honest bottom line

The case for a managed AI agent is not that managed agents are smarter. On a good day, a well-built in-house agent and a well-built managed one do the same work equally well.

The case is that the good day is not the one that decides whether this was a smart investment. Month four decides that: the week the model retires, the month the CRM changes its API, the afternoon an inbound email contains something it should not, the morning a regulator asks a question about a system you have not looked at closely since launch.

If you have someone whose actual job is to be watching on those days, build it in-house and keep the money. If you do not, the recurring fee is not really buying you an agent. It is buying you somebody whose Tuesday gets ruined instead of yours, and a record you can hand to anyone who asks.

_This article describes regulatory requirements in general terms and is not legal advice. Your obligations depend on your licenses, your state, and the specifics of your practice. Talk to your own counsel or compliance advisor before relying on any of it._

If you want a straight assessment of what an agent in your firm would actually require to run safely, [see how our managed service works](/managed-ai-agents/) or [book a free consultation](/book-a-consultation/). We will tell you honestly if the answer is that you are better off waiting.

## About the author

![Jay Taheri, Founder & CEO](/assets/media/jay-taheri-founder-280x300.jpg)

[Jay Taheri](/about/)

Founder & CEO

Jay Taheri is the founder and CEO of AI Abstraction. He has more than 25 years of experience in system design and engineering, programming, system reliability, and machine learning, with roles including Senior Software Engineer at Bloomberg, Radianz, and systems engineer at Credit Suisse. He holds a certificate in Designing and Building AI Products and Services from MIT Professional Education.

Jay founded AI Abstraction in 2024 to help small and mid-sized businesses put AI to work without the buzzwords and complexity that usually come with it. The company designs, builds, and hosts AI agents that handle real, day-to-day tasks, like answering calls, reading email, and qualifying leads, then stays on after launch to keep every agent monitored, tuned, and running.

When you call, you talk to Jay directly, the person who actually builds and runs your agent, not an account manager reading from a script.

---

## About AI Abstraction

AI Abstraction designs, builds, hosts, and maintains AI agents and workflow automation for small and mid-sized businesses. Based in Ridgewood, NJ, serving the New York–New Jersey metro area and remotely nationwide. Call 551-246-4338 or visit https://ai-abstraction.com/contact/.
