Generative AI is costing Amazon far more than expected. The company has seen a growing number of AI projects exceed their budgets by wide margins due to coding errors and insufficient cost controls.
According to the Financial Times, the issue was discussed during an internal meeting of Amazon’s senior engineers this week. They warned that moving traditional software workflows to large language models is creating unexpectedly high costs that are difficult to predict. The company is already developing automated safeguards designed to prevent similar budget overruns in future AI projects.
One Amazon executive admitted that accurately measuring the true cost of generative AI remains a major challenge.
“It’s hard to figure out how much anything involving AI actually costs,” the executive said.
The most expensive example involved an internal project that used Anthropic’s Claude Sonnet model to match author information with product listings on Amazon’s marketplace.
The project ultimately failed, but not before costing the company roughly $1.8 million. That exceeded its original budget by 860%, and the overspending went unnoticed for nearly five months.
During the meeting, engineers warned that coding mistakes that would be virtually free in traditional software development can become “catastrophically expensive” when large language models are involved. The reason is simple: every request sent to an AI model consumes billable tokens, and those costs can accumulate rapidly.
The Claude project was not an isolated case. One AI-powered financial auditing project exceeded its budget by approximately $541,000.
Another initiative aimed at improving delivery speeds across Amazon’s logistics network generated about $134,000 in unexpected costs before the problem was discovered. It took the company two weeks to detect the overspending.
Amazon acknowledged the incidents but said they are part of the learning process that comes with adopting new technology.
“As with any new technology, we experiment, learn, and improve how we use it, including how we increase cost efficiency,” the company said.
At the same time, Amazon rejected the idea that these incidents reflect a broader pattern.
“Cherry-picking a handful of small, isolated examples where teams are learning from one another and portraying them as standard business practice does not reflect how Amazon teams use AI,” a company spokesperson said.
Amazon has already had to rethink some of its internal AI policies. Earlier this year, it shut down an internal leaderboard based on employee activity in its Kiro developer platform after it began encouraging so-called “token maxing,” where employees deliberately increased AI usage to climb the rankings.
Despite the multimillion-dollar overruns, the costs remain relatively small compared with Amazon’s overall business. The company generates around $180 billion in quarterly revenue and expects to spend roughly $200 billion in capital expenditures this year, much of it on AI infrastructure and data centers.
One factor behind the rising costs is the industry’s shift toward token-based pricing. Anthropic, OpenAI, and other AI providers are increasingly moving enterprise customers away from fixed subscriptions and charging based on the number of tokens processed by their models.
While this pricing model allows for greater flexibility and scalability, it also makes expenses much less predictable, especially when autonomous AI agents repeatedly call large language models.
As a result, more companies are switching to lower-cost AI models or open-weight alternatives that can run on their own infrastructure without paying for every token.
The soaring cost of generative AI has become a growing concern across the technology industry. Some companies have reduced headcount while spending several times more on AI tokens than they previously spent on employee salaries. Uber, for example, reportedly exhausted its annual AI budget in just four months.
Nvidia has observed the same trend.
“For my team, the cost of computing is far greater than the cost of employees,” said Bryan Catanzaro, Nvidia’s Vice President of Applied Deep Learning.