Disclosure: Some links on this page are affiliate links. If you purchase through them, we may earn a commission at no extra cost to you. Full affiliate disclosure.

If your team runs its own AI stack — a gateway in front of a handful of model providers, a retrieval layer, a workflow runner — those three boxes are now a more attractive target than the laptops on your desk. Microsoft has been investigating live intrusions against exactly that pattern, and the findings are not subtle: attackers read container environment variables, planted a hook that captures every API key an administrator saves, and left Monero miners running on the victims' CPUs.
The three pieces of software named are LiteLLM, an LLM gateway; RAGFlow, a retrieval-augmented generation platform; and Kestra, a workflow orchestration tool. All of them are widely self-hosted. All of them were exposed to the internet. And on September 2, CISA added the relevant flaws to its Known Exploited Vulnerabilities catalog, which means this moved out of the research-curiosity column and onto patch lists.
The pattern is worth understanding even if you run none of these tools, because the underlying mistake is the same one that puts small-team admin panels on the open web: something that was meant to be internal got a public IP, and it happened to hold the keys to everything else.
Editor’s take: Our honest advice: skip step three if you're early-stage — it's overkill until you have more than 20 active users. Coming back to it later is faster than doing it twice.
Self-hosted AI infrastructure has quietly become the new exposed-database problem: powerful, useful, and shipped with defaults that assume a trusted network. A gateway sitting in front of model providers is an admin surface with credentials and compute attached, which is exactly what an attacker wants. If it is reachable from the internet and still uses a default or single-factor admin path, treat that as urgent rather than theoretical.
Microsoft assessed with high confidence that the LiteLLM gateway was breached through a chain of two flaws. The first, CVE-2026-42271, is a command injection issue in LiteLLM versions 1.74.2 through 1.83.6, fixed in 1.83.7. The second, CVE-2026-48710, is a Host-header validation bypass in Starlette nicknamed BadHost. Combined in a vulnerable configuration, they let an attacker sidestep API-key authentication entirely and land unauthenticated remote code execution.
What happened next is a tidy illustration of why containers are not a security boundary. The payload read /proc/1/environ — the environment of the container's main process, which in a typical LiteLLM deployment is PID 1 — and went looking for API keys, tokens, passwords and database connection strings. It exfiltrated the results using several different tools, presumably so a blocked outbound route or a missing utility would not kill the operation. Then it pulled down an ELF binary, dropped it in a temp directory, and gave it a name resembling a normal Linux service.
Before settling in, the attackers did the things a careful tenant does: fingerprinted the host, checked open ports, and killed any competing miners already running. Then they used the stolen PostgreSQL connection details to query LiteLLM's backend tables, which hold proxy virtual keys, model configurations and provider endpoints.
That last step is the one I would underline. LiteLLM's job is to centralize your provider credentials — OpenAI, Azure, Anthropic, Gemini, all in one place, which is precisely why people run it. Convenience and blast radius are the same property here. One compromised gateway hands over multiple production keys at once.
The RAGFlow intrusion was slower and, to my reading, nastier. Microsoft saw possible server-side request forgery reconnaissance first, then code execution several days later, and then a hidden Python hook inserted into the application's LLM configuration path.
The hook sat in the TenantLLM credential configuration flow. Every time an administrator configured or saved a provider, it silently captured the API key, model name, provider type and endpoint. Think about what that means operationally: the normal remediation step for a suspected key leak — rotate the key, paste the new one into the admin panel — hands the attacker the replacement. Worse, the altered startup path meant the hook reloaded whenever the service restarted, so a restart did not clean it up.
If you take one defensive idea from this piece, make it this one. "We rotated our keys" is only true if the machine you typed them into is clean.
Kestra was hit through CVE-2026-49869, a critical authentication bypass. The auth logic failed to properly validate paths ending in /configs, which let an unauthenticated attacker create and execute workflows. Kestra ships plugins for shell and Python execution, so "create a workflow" is functionally "run commands."
The attackers made a workflow that ran shell commands in the worker container, inspected the Docker socket and container environment variables, then downloaded and ran XMRig to mine Monero on the victim's CPU.
Across all three incidents, Microsoft recorded the same persistence toolkit: service-account SSH key changes, cron manipulation, hidden temporary relays, processes renamed to look like legitimate services, restart loops, and immutable file attributes so cleanup scripts could not delete the payloads.
On September 2, CISA added seven actively exploited vulnerabilities to its KEV catalog. Alongside LiteLLM, Kestra and Starlette, the list included JFrog Artifactory, Sangoma Switchvox and two SonicWall SMA1000 flaws. Under Binding Operational Directive 26-04, federal civilian agencies had until September 5 for most of them, and until September 16, 2026 for LiteLLM (CVE-2026-59822) and Starlette (CVE-2026-48710).
Two details worth noting. First, CISA has not attributed these to a single actor or campaign — this is opportunistic scanning, not a targeted operation, which is arguably worse for anyone with an exposed instance. Second, Wiz separately reported sustained attacks across 90 days of AI-infrastructure honeypot telemetry, including activity against LiteLLM and exposed MCP services. Nobody has to be specifically hunting you for this to land on you.
The LiteLLM entry in the KEV catalog is a different flaw from the one Microsoft saw exploited: CVE-2026-59822 lets an unauthenticated attacker use a fabricated Bearer token to open an authenticated MCP session and potentially reach configured tools. It affects versions before 1.84.0 and is fixed in 1.84.0. If you patched to 1.83.7 after the Microsoft report, you are not done.
Microsoft's recommendations are unusually concrete, and none of them require a security team:
/proc/1/environ is trivial once someone is in the container. Microsoft recommends a managed secrets system, with separate virtual keys and spending limits per team so one leak has a ceiling.If you need remote admin access from outside the office: putting the panel behind a VPN is the cheapest fix available today. It takes an exposed management port off the public internet without re-architecting anything, and it is the single control that would have prevented all three incidents above.
If you must reach admin panels remotely: put them behind a VPN rather than a public login page. It is the cheapest way to take a management interface off the open internet, and it closes the exact exposure every incident in this report started from.
If your exposure is broader than one gateway, our cloud security guide covers key and credential hygiene in depth, and the small business security checklist is the right starting point when there is no dedicated security function.
There is a temptation to file this under "AI is dangerous." It is not. Nothing here required a novel technique or an AI-specific attack; it is unpatched public-facing software holding good credentials, which has been the top cause of intrusion for twenty years. What changed is the concentration of value. An AI gateway is not just another app server — it is the junction point between your users, your models, your databases and your cloud, and it stores the credentials for all of them in one readable place.
We looked at the offensive side of this shift in our analysis of the Unit 42 agentic breach, where AI compressed a two-week intrusion into ten hours. This is the defensive half of the same story: the tools teams adopted to run AI quickly became infrastructure, and most of them are being administered like side projects. They are not side projects anymore, and the patch deadline on the calendar says so.
This summary is based on published threat reporting about exposed AI infrastructure, not on any scanning we performed.
An AI gateway such as LiteLLM sits between your applications and model providers, and by design it holds the API keys for OpenAI, Azure, Anthropic and Gemini in one place. That concentration makes it a single point of failure: one compromise can expose provider keys, database credentials and the container environment at the same time.
Microsoft documented intrusions against three components: LiteLLM, an LLM gateway; RAGFlow, a retrieval and RAG platform; and Kestra, a workflow orchestration tool. All three were internet-exposed. The attackers stole credentials, established persistence and deployed XMRig to mine Monero.
Yes. CISA added seven actively exploited flaws to its Known Exploited Vulnerabilities catalog on September 2, 2026, including LiteLLM CVE-2026-59822, Kestra CVE-2026-49869 and Starlette CVE-2026-48710. Federal agencies must patch LiteLLM and Starlette by September 16, 2026.
If your gateway was ever reachable from the public internet without authentication, yes. Microsoft's guidance is to rotate any provider key connected to a reachable gateway, move secrets out of environment variables into a managed secrets system, and issue separate virtual keys with spending limits per team.
Look for shell or Python processes spawned by your AI components, especially where that coincides with secret access or writes to application files. In the RAGFlow case, attackers planted a Python hook in the LLM configuration path that reloaded on restart and captured API keys every time an administrator saved a provider setting.
