16 September 2026
Heard In AI

Source podcast

Training Data

Conversations with the people building and researching artificial intelligence.

Aaron Levie welcomes AI-written code. AI-written board decks "kill" him

On Sequoia's Training Data podcast, Box's Aaron Levie explains why generated code feels acceptable while a generated board deck does not: a presentation is still read as evidence of what its author knows and can execute. He admits the double standard — he uses AI for his own brainstorms and decisions — and describes reading posts twice, once for the substance and once to guess who wrote them.

5 min read

Box inspects its heaviest AI users to find workflows worth teaching

Box CEO Aaron Levie says the company keeps a list of who burns the most tokens — not to encourage more spending, but to check whether the usage is waste or a practice worth demonstrating to everyone else. He describes pulling a team into a room within six hours to watch one colleague work, reports two-to-threefold gains in delivered customer-facing functionality in parts of the stack, and explains why Box will not drop code review.

6 min read

Aaron Levie's question for AI memory: what belongs in the weights?

On Training Data, Box CEO Aaron Levie was asked where enterprise AI memory is heading — retrieval, or models whose weights absorb a company's knowledge. His answer started with a lawyer who can see five matters and whose access changes daily, and ended with a wish for a rubric deciding what gets baked in and what stays a lookup.

6 min read

Box's Aaron Levie expects open-weight tokens and closed-model revenue to grow together

On Training Data, Box CEO Aaron Levie describes how his customers actually pick models: a default for asking questions of their files, and hard-nosed accuracy evaluations for the high-volume extraction work where most tokens are spent. He endorses Decagon founder Jesse Zhang's argument that mature workflows migrate to open-weight models, and explains why the big labs' revenue and open-weight token volume can climb at the same time.

7 min read

Box's two rules for software in the agent era: beat the generic agent, then let it in

On Sequoia's Training Data podcast, Box CEO Aaron Levie said any company sitting on customers' data now has two obligations: build an agent measurably better than an off-the-shelf one at its own workflows, and expose the same capabilities to outside assistants like Claude and ChatGPT. He described the tuned search-and-retrieval harness behind Box's agent, the evaluations that track model progress, and his bet that within five years roughly 90% of enterprise tokens will be spent on work nobody asked for directly.

8 min read

Coding was the easy case: Aaron Levie on the slow spread of AI at work

On Sequoia's Training Data podcast, Box chief executive Aaron Levie explains why AI swept through software engineering and is moving far more slowly through legal work, sales and the rest of knowledge work: code is text, engineers fix their own broken connections, and their work already lives in GitHub. His conclusion is that the tedious work of getting AI into other people's workflows — not the models themselves — is where he is betting the money is.

7 min read

Peregrine counts its field engineers as R&D, not a cost center

An engineer who faked a missing editing feature using comment fields told Peregrine what to build next. A hurricane simulator stayed with one city. Co-founders Nick Noone and Ben Rudolph describe how they decide which piece of field improvisation becomes a product — and what they say it now costs to serve a city this way.

8 min read

Before the prompt: Peregrine says agents write about 90% of its integration notebooks

On the Training Data podcast, Peregrine's Ben Rudolph described integration agents that run for hours, inspect a customer's databases and split work among sub-agents, writing roughly 90% of the Python notebooks the company uses to connect public-safety records, under the deployment team's oversight. The conversation put the share of effort that happens before a user types a question at 95% — the preparation that let a Florida county ask why it had suddenly run more than a hundred water rescues.

5 min read

Peregrine's pitch: make police data useful without owning it

On the Training Data podcast, Peregrine founders Nick Noone and Ben Rudolph argue that the public-safety software business has grown by collecting ever more data, and that their company inverts it: join the records an agency already holds, leave ownership with the agency, and lock down who may look. The same logic leads Noone to refuse a company-wide ban on facial recognition, leaving that decision to customers, law and local norms.

7 min read

Peregrine tested its first agent on a case detectives had already finished

On the Training Data podcast, Peregrine co-founder Ben Rudolph describes building the company's first operational AI agent with a police customer that had worked a case ending in the exoneration of a wrongly convicted man, then asked whether an agent could reproduce the same findings. He says the agent, which runs for 30 to 60 minutes over hundreds of gigabytes of case evidence, is now used in a few US departments, including a Wisconsin county where a handful of phone records helped place a suspect. Co-founder Nick Noone says the company deliberately lets customers take the credit.

4 min read

Parag Agrawal expects a web that calls the agent when something changes

On Training Data, Parallel Web Systems founder Parag Agrawal traces how agents multiply web searches — from a weekly credit-risk review across 10,000 small businesses to the meeting-prep agents that run hundreds of searches before his own calls — and forecasts a web that, in a couple of years, tells agents when something worth acting on has changed.

5 min read

Parallel sold patient search agents before it could afford a web index

On Training Data, Parag Agrawal explains how his company Parallel entered web search without first building a giant index: it launched a search agent that crawled after a request arrived, replaced outsourced human data collection for insurance, sales and finance customers, and treated the index as a latency optimization to be grown later. He describes the agent-specific architecture behind it, the 200-millisecond Turbo mode Parallel announced in July, and a Google Cloud deal that puts Parallel Search beside Google Search as a grounding option.

7 min read