Good Morning. For months we've been hearing about "token limits," or barriers that companies put in place to ensure individual employees aren't running up massive AI bills on pointless queries.
But every time it comes up, tech leaders always seem to be hedging: assuring us that limits are high, adjustable and there's always a process for increasing them. Which made me wonder: what's the point of having these limits at all?
It turns out, some CIOs are as confused as we are. They're struggling to figure out what the limits should be and struggling to know whether it's a good thing (great ROI) or bad thing (wasteful usage) when someone hits that limit. In today's story, I'm diving into how a handful of them are navigating this tension between wanting to drive AI adoption while keeping skyrocketing token costs in check.
One of my sources in the story was Sastry Durvasula, chief operating officer at financial services firm TIAA. I was fascinated by what he had to say about why token limits are something of a necessary evil, why he doesn't always approve requests to increase them, and why he sees them disappearing in the future. Below I'm bringing you some additional, exclusive insights from our conversation.
WSJ Leadership Institute: Why do you feel like hard token limits are necessary?
Durvasula: We've already seen the horror stories of tokenmaxxing and companies burning through their budgets. There's obviously a balancing act. You have to get employees to start using AI, and then you have the real question of whether the juice is worth the squeeze from an ROI perspective.
WSJLI: How do you decide what the token limit should be?
Durvasula: Based on workforce and workload. We have hypothesis on how much token consumption we think a coder would need that does these types of activities in software engineering vs. how much we need for a communicator or a marketer, or a data analyst in our finance team, or an actuarial scientist that is doing some complex modeling.
As we learn more based on their usage, we're adjusting these limits as well. But that's how we've taken the initial stab. And we're probably in version three or four at this point in terms of revising the token limits.
WSJLI: Are they going up or down?
Durvasula: I haven't seen anything going down to be honest yet.
WSJLI: Are the token limits effective at getting employees to think more critically about their AI queries?
Durvasula: [Employees] know that there's a limit and if they cross the limit then they'll have to go and ask their manager or come to the central team for more tokens, which is a painful process. So the rationing helps for them self-govern a little bit.
WSJLI: How often are people going through that process? And how often are they getting approval for more tokens?
Durvasula: The request volume is not that substantial, which is a good thing. The approval rate depends. In some divisions I think the approval rate is definitely more than 50% and in others it's probably like 10%.
WSJLI: How do you decide whether to grant them more tokens or not?
Durvasula: I think we're still in the learning phase to be honest. Like, how do you identify a really good super user? Some of them are burning a lot of tokens, but they're really good use cases.
Right now there's more bureaucracy in the process today... asking them for justification of why they're using this and like what's the business value they're driving, etc. But I think with time, we'll be able to confidently say that these are the super users in our firm and they will have a different level of access to the compute. But I don't think we are there yet.
WSJLI: Are token limits here to stay?
Durvasula: In the long run I don't think it's the right thing to do. There may be a user that has a real killer query or prompt that's going to change the world for the company in that specific function that may be restricted because of the tokens and there may be other queries that are actually not driving value for the firm.
WSJLI: So what's the long term play for keeping AI budgets in check?
Durvasula: Open telemetry and the routing intelligence. Do you really need the big heavy frontier model for summarizing minutes from our meeting today or can you just use a small, lightweight language model just to do the same task?
The AI Safety Debate Is Getting Louder
Silicon Valley's AI safety debate is getting louder, with more of the very people building the technology now saying it could wipe out humans, the WSJ reports.
Anthropic researcher Jacob Coxon resigned this week, warning that his former employer and OpenAI are racing toward tech that could kill everyone within a decade-a view Anthropic's own Evan Hubinger echoed, putting 10-year extinction risk above 10%.
Fueling the alarm are recent leaps toward AI capable of "recursive self-improvement," plus a string of real incidents, like a swarm of advanced AI agents inside OpenAI hacking into the AI platform Hugging Face and covering their tracks afterward.
The WSJ's Sam Schechner spotlights some of the specific doomsday scenarios keeping AI insiders up at night.
Loss-of-control (extinction). Highly capable AI systems that pursue their own goals in ways that conflict with human survival.
Misuse: A bad actor directs a capable AI toward catastrophe, such as engineering a novel bioweapon or manipulating nuclear powers into conflict.
Sub-extinction catastrophes: Massive AI-enabled cyberattacks that cripple power grids or financial systems, triggering broader social breakdown.
The "WALL-E" scenario. A slow slide where humans cede control to machines and lose the ability to steer their own future.
Time for a palate cleanser.
It Folds!
Apple Wednesday showcased both its first foldable iPhone and CEO John Ternus , who before becoming Apple's third CEO last week led the hardware engineering unit behind the foldable feat.
"We've been imagining quite a lot," Ternus said to the hundreds of gathered media in the audience. "We think you're going to like what you're going to see."
The iPhone Duo, which will start at $1,999, enables iPhone users to have two apps open at the same time. The WSJ reports that one highlight of the new device is that the "crease"-where the internal screen folds-is effectively invisible when open, unlike competing Android devices.
On Our Radar
AI researcher Andrew Tulloch made headlines last year when Meta Platforms offered him a pay package worth more than $1 billion. And now he is leaving the company, people familiar with the situation tell the WSJ. Tulloch hasn't responded to requests for comment, and Meta declined to comment on his exit. Tulloch previously had spent 11 years at Meta before leaving in 2023 for OpenAI and then co-founding Thinking Machines Lab.
The rising class of newly minted millionaires in the San Francisco Bay Area are just like you and me... kind of. "They are not flashy," Garret Spiecker, a senior managing director at Citizens Private Bank, tells the WSJ. But the limiting factor is not an aversion to flying private or buying expensive watches, say, but having enough time to indulge, with the intra SF-AI race demanding all their waking hours. The most commonly cited splurge according to one former OpenAI employee? An espresso machine.
Deere is rolling out a conversational AI assistant called JD to provide farmers quick answers to queries about the best time to plant or harvest crops and other business decisions. The WSJ reports that the Illinois-based equipment maker is counting on JD, available for free, to drive customers toward buying more Deere equipment and technology, including software for their harvesters, crop sprayers, planters and other equipment.
Days after Mistral completed a funding round led by Samsung Electronics that lifted its valuation above $24 billion, the French AI company announced plans to partner with Cloudera to meet growing demand for sovereign AI, the WSJ reports. Mistral will integrate its models with Cloudera's hybrid data platform over the coming weeks so that enterprises can deploy and train them with proprietary data within controlled environments.
The WSJ Technology Council
The WSJ Tech Council brings together CIOs, CTOs and CISOs advancing innovation and shaping the future. Join this trusted community where tech executives connect with peers to explore emerging trends and gain the perspective they need to stay ahead of disruption.
Request Information
About Us
Follow Isabelle Bousquette on LinkedIn, Instagram, X, and TikTok for more behind the scenes on her tech and AI coverage, and lately, her contributions to the WSJ Leadership Institute's new Executive Resilience series, where she's profiling America's top execs about their fitness and wellness habits.
Follow Belle Lin on LinkedIn and X for her latest reporting on enterprise technology and AI.
Steven Rosenbush is chief of the enterprise technology bureau at the WSJ Leadership Institute. He also has a column. You can follow him on LinkedIn.
Tom Loftus is the editor of The Morning Download. He suggests following Isabelle, Belle and Steve on their various social channels. But if you insist, here's his LinkedIn.