Skip to content
back to the archive page
#Management

Every: When It Is Time to Turn an AI Experiment into a Product (Category #Management)

When a new AI model comes out, it is easy to collect enough ideas to rebuild the entire product. Meanwhile, the team still has customers, commitments, and a plan. Dan Shipper, co-founder and CEO of Every, suggests creating a small lab that explores new possibilities while the product team develops what already works.

I have already covered his talk about how Every is organized. In this new 18-minute talk at the Lenny and Friends Summit, the interesting next step is how to tell whether an experiment deserves further investment. The company's scale is worth keeping in mind here. According to Shipper, Every had about 30 people when the talk was recorded on September 10, 2026. His proposal for a one- or two-person lab is based on the experience of a small organization.

Shipper separates two modes of work:

  1. Research requires trying different approaches, allowing duplicate experiments, and being comfortable throwing results away.
  2. Product development requires focus, reliability, and a coherent customer experience. When everyone constantly does both, the team pulls itself in different directions.

For the lab, he proposes a “pirate + architect” pair:

  • The pirate quickly builds rough prototypes and searches for something useful.
  • The architect turns a promising discovery into a system that can be maintained and developed. Shipper suggests accepting in advance that about 90% of experiments will end up in the bin. This is his operating assumption for exploration, not a measured success-rate benchmark.

Then comes the most substantial part: people must return to the tool for their work. First, it is tried inside the team or with a few early customers. Then the team checks whether it remains useful after the initial excitement has passed. Shipper explicitly suggests looking at the result a month later.

He names three criteria for moving an experiment forward:

  • Do people use it and return to it?
  • How much better is it than the existing solution? His benchmark is “10×,” but the talk does not give a rigorous measurement method.
  • Can customers be served at an acceptable cost if usage grows?

A good example is Kate Bench, an assistant for editorial revision. Shipper says he spent several years experimenting with AI based on edits made by editor-in-chief Kate. When Kate began using the result and the tool spread among colleagues, there was a reason to bring in an architect.

The architect added tracking for accepted suggestions and for the edits Kate still had to make after the agent. According to Shipper, the amount of work she spent on those edits fell by 12% compared with the previous month. This is an internal Every result, and he does not explain the calculation method in detail. But the sequence itself is useful: working prototype → regular use → investment in the system and measurement → customer validation.

There is a boundary here that is easy to miss. Shipper does not set a threshold for how many users must return, how often, or over what period. In his framework, regular use is a reason to continue development. Commercial viability still has to be tested.

Every's experience should also be transferred with its business model in mind. The company turns even failed experiments into publications that attract readers and potential customers. That way of paying for exploration may not work in another company.

From this approach, I would take one question into a weekly experiment review: who is still using the result in their work a month later, and what exactly does it give them? The answer helps decide where to direct the next week of development.

#Management #AI #Product #Engineering #Metrics

Open video on YouTube