Asteris Logo

Cheap Tokens, Expensive Consequences

News
WIAISERIESWeek in AITECHNOLOGY11th July
Cheaper model access does not make dependable AI cheap. This week's launches, infrastructure deals and research papers show that the larger costs now sit in power, workflow design, evidence, security and human accountability.

AI tokens are getting cheaper, but using them responsibly is not. The larger costs now sit in infrastructure, workflow failures, verification, security and human oversight, which means lower model prices can make poorly designed automation easier to scale.

OpenAI put a $1 price on one million input tokens this week. Around that announcement sat billion-dollar data centres, new chips, government restrictions, workplace bans and research into why agents forget, misuse evidence and act at the wrong time. The price of intelligence is falling. The price of getting it wrong is not.

The bill beneath the benchmark

OpenAI launched GPT-5.6 in three tiers, with the cheapest priced at $1 per million input tokens and the flagship aimed at stronger reasoning, coding and computer use.1 That is the familiar story of generative AI: capability improves while the unit price falls. The less familiar story is what has to be built, financed and powered before that cheap token can appear on a customer's screen. The model may be getting cheaper to call, but the system required to serve it is becoming more capital intensive.

Meta plans to spend C$13 billion on a data centre in Alberta that could draw up to 1.8 gigawatts of power.2 Bank of America has extended OpenAI a $520 million credit line, its first reported loan, as the company moves towards a more mature capital structure.3 SK Hynix has pursued a major US listing while demand for high-bandwidth memory continues to rise, and Meta is preparing an in-house chip while still buying heavily from Nvidia.45 These are not side stories to the model race. They are the model race in physical form.

The important change is not that AI uses a lot of electricity or expensive chips. That has been true for years. The change is that access to power, memory, financing and long-term supply is becoming a strategic capability in its own right. A model advantage can disappear after the next release, while a power contract, chip programme or data-centre lease can shape costs for years.

That creates a strange market. The most visible layer looks increasingly interchangeable, while the least visible layers become harder to reproduce. A startup may gain access to the same model as a global platform, but it does not gain the same inference economics, purchasing power, distribution or tolerance for losses. Equal access to an API does not create equal access to the business behind it.

This matters to smaller companies because falling model prices can create false confidence. The cost of one prompt may be trivial, yet the cost of embedding an agent across customer service, reporting, research or content operations can rise quickly once reliability, monitoring, storage, security and human review are included. The practical question is no longer whether a model is affordable in isolation. It is whether the complete workflow produces enough value to justify the compute and supervision it consumes.

From answers to assigned work

GPT-5.6 was launched alongside ChatGPT Work, a product designed to gather information across applications and files, divide a goal into smaller tasks and produce finished documents, spreadsheets, presentations and web applications.6 That pairing matters more than the model announcement alone. It turns the model from an answer engine into a worker inside a defined environment. The product claim is not simply that it knows more, but that it can carry responsibility further through the task.

Other launches pointed in the same direction. OpenAI introduced voice models that can listen and speak at the same time, SpaceXAI pushed Grok 4.5 around coding and agentic work, and Mistral moved into physical navigation with a single-camera robotics model. The common thread is not a new interface style. It is the attempt to place AI closer to the point where work is actually performed.

That shift exposes a problem hidden by chat. A poor answer in a conversation can be ignored, corrected or regenerated. A poor action inside a workflow can update the wrong file, send the wrong message, cite the wrong source or move a task forward before anyone notices. The moment AI stops advising and starts acting, product quality becomes inseparable from permissions, recovery paths and review.

The winners will not necessarily be the organisations using the most advanced model. They will be the ones that have made a careful decision about the unit of work they are willing to delegate. That means defining the input, the expected evidence, the acceptable range of action, the point of human approval and the procedure for reversing a mistake. A vague instruction produces a vague agent, however capable the underlying model may be.

This is also why the claim that AI will simply replace jobs remains too blunt to be useful. The Microsoft cuts reported this week sit beside spending on data centres, evaluation, security and model operations because automation moves work before it removes it. Some tasks shrink, while new work appears around system design, checking, procurement, policy and exception handling. The human contribution becomes less about producing every intermediate step and more about deciding what the system is allowed to do.

What does an AI content tool actually do?

An AI content tool should do more than generate text on command. It should gather relevant business context, preserve the brand's point of view, turn existing material into usable formats, show the source of important claims and keep a person in control of what is published. Anything less is a faster blank page, not a dependable content system.

That distinction matters for anyone asking how to automate Instagram content creation. The weak version produces a month of captions from a generic prompt, then leaves the owner to fix repetition, tone and factual mistakes. The stronger version starts with the business's own photographs, offers, products, services and previous posts, then helps organise that material into a coherent schedule. The value comes from reducing unfinished hand-offs without erasing the judgement that makes the business recognisable.

This is the standard AI content tools should be held to. A useful system should help a restaurant turn a menu change into several distinct posts, help a salon reuse evidence of its work without repeating the same caption and help a product brand connect images to a planned campaign. It should not flatten all three into the same polished language. Good automation protects specificity instead of sanding it away.

That is also the practical argument behind AI content generation for small business. Small teams do not need a miniature media department pretending to be autonomous. They need a system that can carry repetitive parts of the work while leaving pricing, claims, taste and final approval with the owner. Instagram AI content becomes useful when it reduces the operational burden without inventing a personality the business never had.

The same principle applies beyond marketing. A coding agent needs the repository's conventions and access boundaries. A research agent needs source requirements and a stopping rule. A reporting agent needs definitions for the numbers it is allowed to combine. Context is not decorative input added after the model is chosen. It is the structure that makes delegation possible.

Trust becomes part of the product

Alibaba's reported ban on Claude Code shows how quickly a capable tool can become unusable inside a company if its data behaviour cannot be explained to the security team.7 CISA's use of Anthropic's Mythos model to inspect government code shows the opposite side of the same issue: AI can be trusted with sensitive work when the use case, controls and review process are explicit.8 The question is not whether the model is generally safe or unsafe. It is whether the organisation can describe the trust boundary for this particular task.

The delayed rollout of GPT-5.6 after US national-security review made that boundary visible at government level.9 Frontier models are now capable enough that a release decision can be treated as a security event rather than a normal product launch. That does not mean every delay is wise or every restriction is proportionate. It means the old assumption that software should ship first and be governed later is becoming harder to defend.

The same tension appeared in proposals for AI-run companies in Argentina, calls for broader oversight of general-purpose models in UK financial services and warnings that international rules are falling behind deployment. None of these stories removes human liability. In fact, they underline it. A company can automate decisions, but it cannot automate away the need for someone to answer when those decisions cause harm.

This is where the phrase human in the loop starts to lose meaning. A person technically present in the process may still lack the information, authority or time to intervene. Real oversight requires access to the evidence, a clear right to stop the system and enough understanding to recognise when the output is wrong. Otherwise, the human is not a control. They are a signature added after the decision has already been made.

Trust therefore becomes a product feature rather than a policy page. Buyers will ask where the model runs, what data it can see, how actions are logged, which sources support the answer and what happens when access changes across jurisdictions. A vendor that cannot answer those questions may still have impressive benchmarks. It will struggle to become part of serious work.

Research moves to the surrounding system

The research published this week reinforces the same shift. ReContext argues that long-context models often fail to use relevant evidence already present in the prompt, then improves performance by replaying query-relevant material before final generation.10 The result challenges a common assumption that more context automatically produces better reasoning. Information can be available to a model without becoming usable evidence.

Memory papers such as AutoMem make a related point. Long-horizon agents do not only need larger context windows; they need methods for deciding what to keep, what to discard and what can be reused later.11 A transcript is not a memory system, and a folder of documents is not organisational knowledge either. Useful memory requires selection, structure and judgement about future relevance.

Evaluation has the same problem. A paper auditing LLM-as-judge methods found that apparent performance can change when the judge changes, which means some model improvements may partly reflect the evaluator rather than the system being evaluated.12 That is not a minor measurement issue. If an organisation uses one model to grade another, the quality of the judge becomes part of the product claim.

Other papers this week examined proactive agents, citation verification, online safety monitoring and multi-agent research. They ask different technical questions, but the direction is consistent. Researchers are spending more attention on when an agent should act, whether its sources can be checked, how it should manage evidence and whether a person can reconstruct the path to a result. The benchmark is moving from task completion towards accountable task completion.

This is a healthier direction for the field. Capability research expands what systems can attempt, while surrounding-system research determines whether those attempts can be trusted outside a demo. Both are necessary, but the second category has been underpriced in product decisions. A higher score is easy to announce. A reliable evidence trail, sensible memory policy and credible evaluation process are harder to show in a launch video.

The cheap part is misleading

Cheaper tokens are real progress. They widen access, make experimentation less risky and allow smaller teams to use capabilities that once belonged only to well-funded companies. But lower input prices do not make the full system cheap. They make it easier to put an imperfect system into more places.

The resulting costs rarely appear on an API pricing page. They arrive as incorrect actions, insecure data flows, fabricated sources, unreviewed content and employees asked to supervise more automation than they can meaningfully inspect. Cheap inference can magnify expensive organisational weaknesses. The effect is not inevitable, but avoiding it requires deliberate design.

For founders, marketers and small business owners, this should change the buying question. Do not ask only which model is smartest or which tool produces the fastest first draft. Ask what the system knows about your work, what evidence it can show, what it is allowed to change and who retains the final decision. Those questions reveal more about the eventual cost than the token price does.

The next model will probably be cheaper again. The consequences will depend on what we connect it to, what we permit it to do and whether a person still has enough visibility to intervene.

Sources

Footnotes

1

GPT-5.6 model launch and pricing, OpenAI

2

Meta's planned Alberta data centre investment, Reuters

3

Bank of America's first reported loan to OpenAI, Reuters

4

Meta's in-house AI chip production plans, Reuters

5

SK Hynix's US market debut and AI memory demand, Reuters

6

ChatGPT Work product announcement, OpenAI

7

Alibaba's reported Claude Code workplace ban, Reuters

8

CISA's use of Anthropic's Mythos for code auditing, Reuters

9

US approval for the broader GPT-5.6 rollout, Reuters

10

ReContext research on evidence use in long-context models, arXiv

11

AutoMem research on trainable memory management, arXiv

12

Audit of reliability in LLM-as-judge evaluation, arXiv