LLMjacking is the abuse of stolen or leaked cloud credentials to run large language model inference at the victim’s expense. The attacker uses your account’s access to services such as Amazon Bedrock, Azure OpenAI or Google’s Gemini API, and you pay for every token. The term was coined by the Sysdig Threat Research Team in May 2024.
This guide covers the part most definitions skip: the exact log events that reveal LLMjacking on AWS, Azure and GCP, and which of those events you can see without changing any settings.
How a credential ends up being used this way
LLMjacking almost never starts with a sophisticated break-in. It starts with a credential that was easy to create and easy to lose.
The common route runs through AI pilots. A team wants to test a model, so someone creates an access key for a notebook, a local script or an MCP server configuration. The key gets a broad policy so that nothing breaks during the pilot. It is then copied into a repository, a shared notebook, a .env file or an mcp.json file. Automated scanners watch public repositories for cloud credentials, and a leaked key can be found and tested quickly after it is pushed.
The other route is an ordinary compromise. In the case Sysdig first documented, the attackers exploited a vulnerable web application and took the cloud credentials stored on it. The credentials were not created for AI at all; they simply had permission to call it.
On AWS, Bedrock adds its own credential type. Amazon Bedrock API keys come in two forms:
- Short-term keys last up to 12 hours or the length of the session that generated them, whichever is shorter. They inherit the permissions of the principal that created them and work only in the region where they were generated.
- Long-term keys create an IAM user behind the scenes, attach the policies you select, and last until the expiry you choose. The creation shows in CloudTrail as CreateServiceSpecificCredential.
AWS recommends long-term keys only for exploring Bedrock, and says to move to short-term credentials for anything with real security requirements. That advice is sound, and it describes the exact trap. Pilots are exploration, so they get long-term keys. When a pilot turns into a production service, nobody goes back to swap the key.
Classic IAM user access keys, which start with AKIA, carry the same risk. They stay valid until someone rotates or deletes them.
The LLMjacking attack sequence in CloudTrail
Once an attacker has a working AWS credential, the activity follows a recognisable order. Each stage leaves events in CloudTrail.
| Stage | What the attacker is doing | Events you would see | Why it stands out |
| 1. Validate the key | Checking the key works and which identity it belongs to | sts:GetCallerIdentity | First use from an IP address or network (ASN) never seen for this identity |
| 2. See what models exist | Listing available models, often across regions | ListFoundationModels | Rare for a production identity, and often from regions it has never used |
| 3. Check whether you are watching | Reading the account’s model invocation logging settings | GetModelInvocationLoggingConfiguration | Almost no legitimate workload asks this |
| 4. Use the models | Running inference, often at high volume | Bursts of InvokeModel, InvokeModelWithResponseStream, Converse, ConverseStream | Volume and model mix far above the identity’s baseline |
| 5. Look for more access | Working out what else the key can do | ListAttachedUserPolicies, ListRoles, GetAccountAuthorizationDetails, SimulatePrincipalPolicy, often with bursts of AccessDenied | A new API family for an identity that only ever called Bedrock |
Sysdig also observed attackers calling InvokeModel with a deliberately invalid parameter. The call fails validation, which confirms the key has model access without producing a billable response. A run of failed invocation calls from a new source is therefore worth treating as part of stage 1.
The logging check is the most useful signal
GetModelInvocationLoggingConfiguration returns whether the account records prompts and responses, and where it sends them. An attacker calls it to find out whether their usage will be visible. Very few legitimate applications ever need to ask. An identity that checks your logging configuration and then starts invoking models is a high-quality alert with few false positives, and it arrives before most of the cost does.
You do not need data events to see the abuse
According to the AWS documentation on Bedrock and CloudTrail, InvokeModel, InvokeModelWithResponseStream, Converse and ConverseStream are logged as management events. Management events are recorded by default and kept for 90 days in CloudTrail Event history, so the invocation burst is visible in any AWS account without extra configuration.
Data events are a different matter. Agent calls such as InvokeAgent and InvokeInlineAgent, knowledge base calls such as Retrieve and RetrieveAndGenerate, and S3 object reads are data events. They are off by default and cost extra. If the attacker moves from running models to reading data through an agent, you only see it if CloudTrail data events [/blog/cloudtrail-data-events/] are on.
Stage 5 is where LLMjacking stops being a billing problem. An identity that can call AssumeRole into another account can take the attacker from your AI bill to your customer data.
The same attack on Azure and GCP
The pattern is the same on the other two clouds. The places the evidence lands are different, and in two cases there are gaps worth knowing about.
| Stage | Azure | GCP |
| Credential created or retrieved | Microsoft.CognitiveServices/accounts/listKeys/action in the Activity Log when a key is read; Entra audit logs for new app registrations and secrets | google.api.apikeys.v2.ApiKeys.CreateKey or CreateServiceAccountKey in Admin Activity audit logs |
| Credential used from a new place | Key-based calls to Azure OpenAI do not appear in Entra sign-in logs. Service principal use appears in AADServicePrincipalSignInLogs with IP address and credential key ID | authenticationInfo.serviceAccountKeyName names the key and requestMetadata.callerIp shows the source, where the call is logged |
| Model usage | AI Services resource logs, which need a diagnostic setting per resource | Most read and invoke activity is in Data Access audit logs, off by default except for BigQuery |
| Main fix | Set disableLocalAuth so the resource accepts only Entra authentication | Replace API keys and service account keys with attached service accounts; restrict any key that must exist |
Azure: key use is not a sign-in
Key-based access to Azure OpenAI and Azure AI Services bypasses Microsoft Entra ID entirely. Using a key is not an Entra sign-in, so it does not appear in sign-in logs at all. Teams that watch service principal sign-ins for unusual locations will see nothing.
What does appear is the moment someone retrieves the key. The listKeys action is recorded in the Azure Activity Log, which is on by default for 90 days. An unexpected identity reading keys from an AI Services account is the Azure equivalent of the AWS logging check. The lasting fix is to set disableLocalAuth on the resource, which forces every caller through Entra, where the call is tied to an identity and logged.
GCP: API keys identify a project, not a principal
A Google Cloud API key identifies a project. It does not identify a user or service account, so a log entry for a call made with a key tells you which project paid, not who made the call.
Google has tightened this for Gemini. Research from Truffle Security in February 2026 showed that API keys created for other services, such as Maps, could call the Gemini API once it was enabled on the same project. Google’s Gemini API key documentation now states that the Gemini API rejects requests from unrestricted standard keys, blocks unrestricted keys that have been dormant for a long period, and will reject standard keys entirely from September 2026 in favour of a new key type. Check which key types your projects still use.
For service account keys, the audit log records which key was used and the caller’s IP address, but only for calls that are logged. Most Vertex AI and data read activity falls under Data Access audit logs, which you have to enable per service.
Detecting LLMjacking before the bill arrives
Most organisations find LLMjacking on their invoice. Three layers of detection move that discovery earlier.
The first layer is a behaviour baseline per identity. Over 14 to 30 days, record the API families each identity calls, the regions, the source networks and ASNs, user agents, hours of activity, call rate and error rate. A machine identity that has only ever called Bedrock from one region inside your VPC is easy to baseline. When it suddenly calls from a hosting provider’s network in a region you do not use, the deviation is obvious.
The second layer is a small set of specific rules that do not depend on a baseline:
- any call to GetModelInvocationLoggingConfiguration by an identity that is not your security or platform tooling
- ListFoundationModels across several regions in a short window
- first use of IAM or STS discovery APIs by an identity that normally calls only Bedrock
- listKeys on an Azure AI Services account by an identity that has not done it before
- creation of a Bedrock long-term key (CreateServiceSpecificCredential) outside your approved pipeline
The third layer is a cost alert on model spend. On AWS, use Cost Anomaly Detection or AWS Budgets filtered to Amazon Bedrock. On Azure, use cost anomaly alerts and budgets in Cost Management. On GCP, use Cloud Billing budgets and alerts scoped to the Vertex AI and Gemini services. Billing data lags usage, so treat this as a backstop rather than the primary alert.
Expect some false positives. A legitimate new workload, a load test and a genuine model evaluation all produce invocation bursts. Tie those to a change record or a known deployment identity, and tune the rules to ignore them, rather than raising the threshold for everything.
Preventing LLMjacking
Prevention is mostly about removing the long-lived credentials that make the attack cheap.
Use workload identities wherever the code runs inside the cloud. A Lambda function, ECS task, Azure managed identity or GCP attached service account gets temporary credentials that rotate automatically and are never written to a file.
Where a key must exist, give it an expiry and a narrow policy. On AWS, prefer short-term Bedrock API keys over long-term ones, as AWS itself advises. You can deny API key use for an identity with a policy on bedrock:CallWithBearerToken, and AWS publishes example service control policies for Amazon Bedrock for organisation-wide limits.
Limit which models and regions an identity can invoke. An identity that can call only the one model it needs, in one region, is a far smaller prize.
Turn on secret scanning for your repositories, and treat notebook files and MCP configuration files as places secrets end up.
On Azure, set disableLocalAuth on AI Services resources. On GCP, enforce the iam.disableServiceAccountKeyCreation organisation policy where it is not already on, and restrict any remaining API key to the specific API it needs.
If a Bedrock key is compromised, AWS documents how to revoke long-term and short-term keys. Long-term keys can be deactivated or deleted. Short-term keys cannot be revoked individually, so you deny the session with a policy instead.
Common questions about LLMjacking
LLMjacking is the use of stolen or leaked cloud credentials to run large language models on the victim’s account, so the victim pays for the inference. Attackers use the access themselves or resell it.
Mostly from leaks: keys committed to public repositories, left in notebooks or configuration files, or taken from a compromised server. AI pilots are a common source because they often use long-lived keys with broad permissions.
It depends entirely on which models the attacker runs and how much. To estimate your own worst case, take the per-token price of the most expensive model your credentials can reach, multiply it by that model’s tokens-per-minute quota, and repeat for each region where the model is enabled. Your cloud’s pricing page and service quotas console have both numbers.
Yes. Bedrock InvokeModel, InvokeModelWithResponseStream, Converse and ConverseStream calls are management events, so they are recorded by default. The prompts and responses are not; those need Bedrock model invocation logging, which is off by default.
Where a platform helps
The detection logic above can be built with native tools and a SIEM. The hard part is doing it for every machine identity across three clouds, and doing it fast enough to matter.
Cy5’s ion platform approaches this through the identity. It evaluates each control-plane event as it arrives rather than on a scheduled scan, and keeps per-identity behaviour baselines for machine identities as well as people. A key being validated from a new network, a logging-configuration check or a sudden invocation burst shows up as a deviation for that specific identity, scored against what the identity can reach.
To be clear about scope: ion detects the key’s use, not the leak itself. It sees the moment a leaked credential starts making calls in your cloud, which is when the cost and the risk begin. For more on how this fits the wider problem, see our guide to AI agent security and the ion platform overview.
What the logs cannot tell you
Three limitations are worth stating plainly.
Billing alerts are lagging indicators. By the time a cost anomaly fires, the usage has already happened.
Behaviour baselines need history. A brand-new identity has no baseline, and an attacker using it looks no different from its owner until the rules above catch a specific action.
And if Bedrock model invocation logging was never switched on, you cannot go back and prove what was generated on your account. CloudTrail shows that the models were called, how often and by which identity. It does not show the prompts or the output. If what was generated matters to you, for legal or reputational reasons, turn invocation logging on now, before you need it.
For the wider AWS picture, see our guide to Amazon Bedrock security [/blog/aws-bedrock-security/], and for why these credentials deserve their own inventory, our explainer on non-human identities.