The problem.
Notion had voice input on mobile. Ryan Nystrom wanted that capability available through its web and desktop product.
What changed.
He gave Codex the mobile implementation, described the intended experience and provided a way to verify the work.
What was reported.
Nystrom estimated his own work at three or four hours. He reported shipping the feature the next day for user testing. Source: OpenAI
What the evidence can tell us
This is one engineer’s account. The source’s two-week comparison is his estimate of how two engineers might have approached the work earlier; it is not a measured control or a general productivity ratio.
What we take from it.
Extending an existing feature creates a useful starting point: the team already has a working example of the behavior it wants. We would turn that example into explicit acceptance criteria before asking an agent to write code. Those criteria should describe what the user can do, where the experience differs between devices, and how the team will recognize a complete result.
Moving a feature between platforms still requires product decisions. Voice input on the web needs a considered response when microphone access is denied, recording is interrupted or a browser behaves differently. The team should decide how recording state is communicated and how a user can recover. A successful demonstration is a starting point for checking those less convenient paths.
The existing codebase also supplies conventions that a generated implementation needs to respect. We would point the agent toward the relevant components, data boundaries and test patterns, then review its changes against those conventions. Keeping the change focused makes it easier to see whether the implementation solves the intended problem or introduces unrelated behavior that will be expensive to maintain.
The practical measure is how quickly the team reaches a feature people can use and the team can support. That includes review, accessibility checks, testing and the work created by early feedback. We would track those costs alongside implementation time, preserve a clear rollback path and assign an owner for follow-up. The resulting evidence is more useful for the next delivery decision than a dramatic speed comparison.
How to evaluate a similar idea.
Start with your situation and a question you can test. These are evaluation steps we would discuss before choosing an implementation.
- 01
Give the agent a working reference
Identify the existing behavior, relevant files and platform differences. Write acceptance criteria that a reviewer can check independently.
- 02
Design the incomplete states
Specify denied permissions, interrupted input and recovery. Include keyboard operation and clear recording feedback in the review.
- 03
Review a focused change
Check integration, data handling and existing conventions. Run tests that verify user behavior and the important failure paths.
- 04
Learn from a controlled release
Collect early user feedback and record follow-up work. Compare the full delivery effort before applying the approach to a larger feature.
Sources & credits.
- OpenAIWhat Codex unlocks for Notion
Published 9 June 2026 · Checked 14 September 2026
- NotionWhy we built Notion
Publication date not disclosed · Checked 14 September 2026
- Work credited to
- Notion engineer Ryan Nystrom
- Technology / platform
- OpenAI
- Analysis & explanation
- Cactera. Company wordmarks identify the article subjects.
Independent Cactera analysis of publicly documented work. Cactera was not involved in this work. Company names identify the subjects, not Cactera clients or partners.
