A year ago, I treated AI tools as autocomplete with better taste. That is no longer an accurate mental model.

The real shift isn't raw token generation speed or synthetic benchmarks. It is scope. The baseline unit of work got bigger: I used to ask for an isolated function; now I ask for an entire feature, point the agent at the surrounding codebase, and review the resulting diff.

My daily role shifted from typing syntax to specifying constraints and vetting code. While that sounds like a subtle semantic tweak, in practice it rewires how an engineer spends a working day.

The Shift in the Prompt

The transition is best illustrated by how the questions we ask have evolved:

TypeScript

// The old unit of work: isolated, micro-level syntax
"write a debounce function"
// The new unit of work: contextual, diagnosis-to-resolution
"the search input on /experiments feels laggy on mobile — find the root cause and fix it"

In the first case, the developer does the diagnosing, architecture, and wiring—delegating only the keystrokes. In the second case, the tool is tasked with tracing state, identifying performance bottlenecks across files, and proposing an integrated solution.

Where the Model Holds Up

AI excels at tasks with clear boundaries, mechanical predictability, and high boilerplate overhead:

  • Boilerplate and scaffolding: Data models, CRUD endpoints, TypeScript interfaces, and repetitive test suites.

  • Mechanical refactoring: Renaming patterns, migrating component patterns across directories, or updating deprecated library calls.

  • Codebase onboarding: Rapidly breaking down unfamiliar modules, legacy files, or dense algorithms before making edits.

Where the Model Breaks Down

The bottleneck is no longer code syntax; it is taste and context:

  • Product judgment: AI cannot decide what not to build. It generates candidate solutions effortlessly, but cannot discern which compromise fits the long-term roadmap or past technical debt.

  • Information architecture & naming: High-level domain modeling requires deep empathy for team conventions and future maintenance—areas where statistical predictions default to generic tropes.

  • System history: Models answer the prompt in front of them; they do not know the historical failures that led to existing edge cases.

The tools got significantly better at answering questions. They didn't get better at asking the right ones.