In-app reader
8 min read
Related Products Advanced URL FilteringCloud-Delivered Security ServicesCode to Cloud PlatformCortexCortex CloudCortex XDRCortex XSIAMPrisma AIRSUnit 42 AI Security AssessmentUnit 42 Incident Response
By:
Published: August 6, 2026
Categories:
Tags:
Share
It’s three a.m., do you know what your AI agent is doing? Unit 42 has responded to a growing number of AI token jacking cases resulting in staggering financial losses.
The financial loss comes from criminals gaining access to API keys used by legitimate developers for access to popular AI platforms. These keys are known as tokens, and their theft is called token hijacking, or token jacking for short.
The unrelenting frenzy of AI adoption and soaring costs of model access are converging into an irresistible opportunity for cybercriminals. Premium pricing on scarce AI processing power means stolen access via tokens can generate a quick and easy profit for attackers. Complex, patchwork billing management and limitless scaling by default can lead to massive financial losses in short periods.
Good security hygiene, combined with cutting-edge native AI protection tools, can prevent losses before they begin.
Palo Alto Networks customers are better protected through the following products and services:
Cortex XDR and XSIAM
The Unit 42 AI Security Assessment can help empower safe AI use and development.
If you think you might have been compromised or have an urgent matter, contact the Unit 42 Incident Response team.
Related Unit 42 Topics AI, LLM, Supply Chain
Token jacking is a new AI-oriented spin on an old technique of stealing access to computing resources.
Establishing a session in service-based computing typically requires authentication, usually involving a username and password, and sometimes a secondary verification method. Many services allow an authenticated user to then generate keys that programs can use on a user's behalf to establish sessions without going through an interactive login to support automated processes. Within a session, the service provider and user have agreed on a structured way to pay to use their service to achieve a pre-defined objective.
AI — in particular, large language models (LLMs) — typically does not have pre-defined objectives. Users can and do carry on long conversations of widely varying complexity, which can consume enormous amounts of the provider’s computing resources. Automated processes also use LLMs to produce iterative content, which they then further process and return to the LLM with additional, related prompts.
To best support this freeform usage, providers typically break both the input prompt and the output data into small chunks called tokens. Regardless of the objective, billing is then based on how many of these tokens are consumed during the session.
Newer and more complex AI models charge more per token, ostensibly because more resources are required to deliver the output. To avoid interruptions in unpredictable workstreams, many providers do not limit the number of tokens an account can consume, instead tallying usage and billing on a cycle.
If an attacker can steal one of these keys, they may find themselves with unlimited programmatic access to tokens that they can then use themselves or resell to other users. Since billing occurs cyclically, the victim might not even be aware of the theft until the attacker has consumed a massive number of tokens.
To better understand token jacking, we must understand transfer stations. Skyrocketing token costs for frontier AI models and regional usage restrictions have spawned a massive gray market of fly-by-night vendors selling AI computing capacity at a fraction of the retail cost.
Figure 1 below shows an example of these advertisements. These services are commonly called transfer stations.
Third parties acting as intermediaries between official AI providers and end users sell these transfer stations. Many of these advertisements appear on Chinese-language marketplaces like Taobao. They promise access to multiple AI services with seller-issued custom credits that are purchased anonymously. Earlier this year, a researcher named Harshal Singh posted a fascinating deep dive into this world.
A large number of these transfer stations run on just a few open-source software platforms like new-api or one-api, which act as proxy services to official AI APIs. These proxy services handle:
Obfuscation
Rotation and authentication of real credentials
Billing
Model routing
Normalization of prompts
In many cases, users of these transfer station services are developers seeking inexpensive AI access. Other use cases are less benign.
Competing nation-states can use these transfer stations' proxy services to access cutting-edge frontier models to train and refine their own models at a fraction of the cost that AI development normally incurs. Transfer stations require access to legitimate API tokens for the associated AI models. Attackers often steal or hijack these tokens from a variety of legitimate sources.
For transfer stations to be cost-effective, their operators require access to a large pool of discounted legitimate tokens for each frontier AI model offered. Purchasing tokens at full price to simply resell them at a discount isn’t profitable, so many operators turn to stolen credentials.
Attackers can use privileged corporate developer accounts they’ve harvested via information stealers or through phishing campaigns to perform the following activities:
Creating new API keys
Provisioning models
Removing billing limits
Disabling critical usage alerts and logging
These developer accounts are readily available for sale by access brokers on dark web marketplaces.
However, a more direct approach is to steal already provisioned access keys. Attackers can harvest these like they do credentials. They can also mine keys from improperly secured file shares or code repositories.
More recently, attackers have stolen these keys using poisoned, self-propagating npm packages downloaded by unsuspecting developers. Once installed, these packages infect any other code releases the developer builds. They steal credentials and access tokens from each environment along the way, amplifying the impact.
Particularly concerning are npm supply chain attacks like Shai-Hulud and Miasma. Attackers could use the huge number of credentials stolen in these campaigns to fuel transfer stations for years.
The financial impact of token jacking can be catastrophic to organizations. Transfer stations can generate tens of millions of API calls per day, resulting in hundreds of thousands of dollars in usage fees.
We’ve responded to cases where attackers stole inadvertently exposed credentials and integrated them into a transfer station within minutes. This led to nearly a million dollars in charges before discovery and containment.
In some of these cases, we connected massive numbers of malicious API queries to domains hosting the new-api proxy service. Figure 2 shows an example of a transfer station frontend marketplace hosted on an IP address running an instance of new-api and connected to an attack.
Organizations impacted by token jacking have very little recourse to recover funds billed by the AI services for using their API tokens. The cost can derail budgets or even force smaller businesses into bankruptcy.
Even unsuspecting developers trying to use transfer stations for legitimate development risk having their prompts routed to inferior models. Furthermore, developers risk having their sessions monitored and mined for sensitive data that could turn them into future victims.
Organizations can protect themselves against token jacking through various methods.
Implement spending limits for AI usage
Ensure that these limits alert organizations if usage changes drastically from an established baseline
Review all privileged accounts that can be used to provision resources or adjust spending limits
Migrate from long-term access keys to short-term bearer tokens to limit the potential window of damage
Use an AI gateway in combination with a machine authentication platform
This can help ensure that all LLM traffic is tied to a verified and managed machine identity, allowing for real-time monitoring of traffic and usage anomalies
Ensure that compute resources include network boundaries where available
This restricts access to corporate infrastructure, preventing compromised keys from being used in a transfer station scenario
Tightly manage development environments to ensure malicious packages do not enter the development pipeline
AI adoption is accelerating at an unprecedented pace. A mindset of “fail fast and break things” has never been more true — or more risky — than it is today.
This mindset brings with it an opportunity for cybercriminals to target vulnerable organizations through token jacking and to cause staggering losses. While innovation cannot be at the mercy of security, there are ways defenders can manage their risk.
Palo Alto Networks customers are better protected through the following products and services:
The Prisma AIRS AI Gateway helps provide a central control plane to secure and govern enterprise AI traffic. By managing API keys centrally, it removes sensitive credentials from developer environments and build systems. Platform teams can gain full visibility into model usage, agent actions, and token spend across teams. Security teams get integrated guardrails that can enforce access policies, prevent data leaks, and set proactive budget limits.
Idira Agentic Identity Security helps provide a comprehensive identity security solution for discovery, control and governance of agentic identities. It provides a central registry of agents with cryptographically verifiable identities, enforces strong authentication and zero standing privileges for agents and provides comprehensive audit trails of agent actions. It also enables agents to secretly retrieve and use secrets and API tokens just in time thereby reducing the attack surface.
Koi Agentic Endpoint Security helps discover all software on your endpoints, both binary and non-binary, from installed applications to code packages and AI artifacts. From there you can govern it, whether that means removing a risky or malicious item, or holding new package versions back until they've had time to establish a reputation under public scrutiny.
Cortex Cloud, XDR, and XSIAM customers are better protected from the topics discussed within this article with cloud runtime security operations monitoring their continuous integration and continuous development (CI/CD) pipelines to ensure that the latest npm packages integrated into test and production environments are monitoring for and preventing malicious code execution.
Using Cortex Cloud’s Identity Security which includes Cloud Infrastructure Entitlement Management (CIEM), Identity Security Posture Management (ISPM), Data Access Governance (DAG) as well as Identity Threat Detection and Response (ITDR), allows clients to monitor cloud identities which may have been compromised as a result of the techniques discussed in this article. Enabling these features helps protect cloud identities.
Advanced URL Filtering identifies known domains and URLs associated with this activity as malicious.
The Unit 42 AI Security Assessment can help empower safe AI use and development.
If you think you may have been compromised or have an urgent matter, get in touch with the Unit 42 Incident Response team or call:
North America: Toll Free: +1 (866) 486-4842 (866.4.UNIT42)
UK: +44.20.3743.3660
Europe and Middle East: +31.20.299.3130
Asia: +65.6983.8730
Japan: +81.50.1790.0200
Australia: +61.2.4062.7950
India: 000 800 050 45107
South Korea: +82.080.467.8774
Palo Alto Networks has shared these findings with our fellow Cyber Threat Alliance (CTA) members. CTA members use this intelligence to rapidly deploy protections to their customers and to systematically disrupt malicious cyber actors. Learn more about the Cyber Threat Alliance.
Table 1 contains indicators associated with recent token jacking activity.
Indicator Context
Go-http-client/2.0,gzip(gfe) User Agent associated with malicious API calls
3.235.109[.]125 Malicious API calls
116.105.166[.]148 Malicious API calls
172.96.142[.]186 Malicious API calls
38.46.219[.]166 Malicious API calls
38.46.219[.]163 Malicious API calls
38.46.219[.]162 Malicious API calls
23.237.196[.]170 Malicious API calls
15.204.106[.]173 Malicious API calls
104.243.42[.]117 Malicious API calls
198.255.70[.]210 Malicious API calls
47.88.103[.]81 Malicious API calls
47.251.72[.]239 Malicious API calls
117.72.74[.]48 Malicious login (Credential Theft)
207.246.106[.]162 Malicious login (Credential Theft)
23.236.182[.]215 Malicious login (Credential Theft)
95.214.112[.]26 Malicious login (Credential Theft)
amutes[.]com Transfer station infrastructure
abb1[.]life Transfer station infrastructure
Table 1. Indicators of token jacking activity.
How Chinese Sell “Claude” Tokens at 5% Cost While Making Millions (Tutorial) – X
The npm Threat Landscape: Attack Surface and Mitigations - Palo Alto Networks Unit 42
**Back to top
Threat Research Center Next: The Frontier AI Vulnerability Burst: Industrializing Autonomous Zero-Day Discovery in Open-Source Software
The Xcode Assassin Returns: A Deep Dive Into the Latest XCSSET Version
Russian Global Webmail Espionage
The npm Threat Landscape: Attack Surface and Mitigations (Updated July 15)
High Profile Threats August 6, 2026 ChainDrop: Inside a Self-Propagating npm Worm
Threat Research August 4, 2026 Almost Half of Malware Samples Communicate Direct to IP
Threat Research August 3, 2026 Pass the Passkey: A Novel Attack Surface in Passwordless Authentication
Threat Research July 31, 2026 The Xcode Assassin Returns: A Deep Dive Into the Latest XCSSET Version
High Profile Threats July 15, 2026 The npm Threat Landscape: Attack Surface and Mitigations (Updated July 15)
Threat Research July 15, 2026 TuxBot v3: Inside an IoT Botnet Framework With LLM-Assisted Development
Threat Research July 7, 2026 Vidar Stealer Unmasked: Code Signing Abuse, Go Loaders and File Inflation
Threat Research June 30, 2026 Phantom Squatting: AI-Hallucinated Domains as a Software Supply Chain Vector
Threat Research June 25, 2026 CL-STA-1062 Targets Southeast Asian Governments and Critical Infrastructure
[Read now](https://unit4
…(truncated for reading performance)
Discussion
Sign in to join the discussion.