Two years ago an AI coding tool meant autocomplete: you typed, it suggested the rest of the line. In 2026 the tools that developers actually reach for are agents. You describe a task, the agent reads the relevant files, makes a plan, edits the code, runs the tests, reads the failures and fixes them, then hands you a diff. Surveys this year put agent use among professional developers somewhere around a third and climbing fast. This post is about what that shift means in practice, from someone who uses these tools daily on a small codebase and a large one.
The three tools everyone compares
The names that come up in every comparison are Claude Code from Anthropic, Codex from OpenAI, and Cursor. There are others, GitHub Copilot's agent mode, Google's Antigravity, open source options like Cline and OpenCode, but by mid 2026 they all converged on the same shape:
- A command line interface you run inside a repo, which is where the agent has full access to files and shell.
- An editor integration for VS Code and JetBrains so you can see diffs inline.
- Cloud agents that run on a branch in the background and open a pull request when done.
- Support for MCP, so the agent can talk to your database, issue tracker or browser through a standard plug.
The differences are in the underlying model, the pricing, and how much of the harness you can script. Which is best changes every few months when a new model ships, so any ranking you read is already stale. The workflow advice below applies to all of them.
What a session actually looks like
Here is a real task from building this site: "the blog generator writes a sitemap, add lastmod dates per post instead of one date for everything". With autocomplete you would open the file, find the function, and type. With an agent:
> add per post lastmod dates to the sitemap in blog-src/build-blog.mjs,
posts have an optional date field, fall back to the build date
Reading blog-src/build-blog.mjs...
Reading blog-src/posts-json.mjs to check the date field...
Editing build-blog.mjs: sitemap() now uses p.date ?? DATE
Running npm run blog...
wrote 16 posts + index + sitemap
Checking public/sitemap.xml... dates vary per post as expected.
Done. Changed 1 file, 3 lines.
Thirty seconds, and the agent verified its own change by running the build and reading the output. That verification loop is the real difference from autocomplete. The agent can be wrong, but it can also notice it was wrong.
Where it genuinely saves time
- Debugging from a stack trace. Paste the error, the agent finds the file, reads the surrounding code, forms a hypothesis and tests it. Teams report this as their single biggest gain.
- Writing tests for existing code. Tedious for humans, mechanical for an agent that can read the function and run the test runner until it passes.
- Mechanical refactors across many files. Rename a concept, migrate an API, update every call site. The agent does not get bored on file 40.
- Unfamiliar codebases. "Where is the retry logic for outbound webhooks?" gets you the file and an explanation in a minute instead of twenty of grepping.
- Boilerplate with a known shape. A new endpoint that looks like the other twelve endpoints. The agent copies the pattern faithfully.
- Shell and config work you do rarely. Makefiles, nginx configs, GitHub Actions, Dockerfiles. Things you would otherwise look up every time.
Where it wastes time
The same tools have well known failure modes, and the difference between developers who love them and developers who hate them is mostly whether they have learned to avoid these:
- Vague tasks. "Make the app faster" produces a flurry of changes of unknown value. Agents do best with a concrete target and a way to check it.
- Design decisions. An agent will pick an architecture confidently and it will often be the most average one. Decide the shape yourself, then let it fill in.
- Silent scope creep. Asked to fix one bug, the agent also reformats the file, renames two variables and "improves" an unrelated function. Review every diff. Tell it to make minimal changes.
- Passing tests by cheating. If the test is hard to satisfy, an agent may weaken the assertion or special case the input. Read the test changes more carefully than the code changes.
- Hallucinated APIs. Calling a library function that does not exist, especially in less common libraries or recent versions. The build catches it, but only if the agent runs the build.
- Long sessions that drift. After an hour the agent's context is full of dead ends. Start fresh with a summary rather than pushing on.
Working habits that make agents pay off
A few practices turn up in every experienced user's advice:
- Write a project instructions file. All the major tools read a Markdown file at the repo root (CLAUDE.md, AGENTS.md, .cursorrules) on every session. Put in it: how to run the tests, the build command, naming conventions, what not to touch. This is the highest return ten minutes you will spend.
- Give it a way to verify. Tests, a type checker, a linter, a build that fails loudly. An agent with a feedback loop is dramatically more reliable than one without.
- Small commits, always on a branch. Let the agent work, review the diff, commit or discard. Never let it run against uncommitted work you care about.
- Plan before code on anything non trivial. Ask for a plan first, correct it, then say go. Most tools have a plan mode for exactly this.
- Treat output as a pull request from a fast junior colleague. Well read, quick, tireless, occasionally confidently wrong. Review accordingly.
The vibe coding question
"Vibe coding", accepting whatever the agent produces without reading it, works for throwaway prototypes and personal scripts. It is how a lot of people who never programmed before are now shipping small apps, and that is genuinely new. For anything with users, data or money attached it fails in the usual ways: security holes, unmaintainable structure, and a codebase nobody on the team understands. The tools are good enough that the temptation is real. Resist it in proportion to the blast radius.
What it means for the job
The honest answer is that the mechanical part of programming, turning a clear specification into working code, is getting cheap. The parts that are not getting cheap: knowing what to build, deciding how the pieces fit, judging whether the output is correct, and taking responsibility when it is not. Developers who use agents well are shipping more, and the skill that matters most has shifted from typing code to specifying, reviewing and testing it. That is a change in emphasis, not a replacement. Reading code carefully has never been more valuable.
Getting started
Pick any of the three, install the CLI, open a repo you know well and give it a small, concrete task with a test you can run. Write the instructions file after the first session, once you have seen what it gets wrong. Within a week you will have a feel for where it helps you. If you want to understand what the model underneath is actually doing, and why it sometimes invents a function that does not exist, our AI basics series starts from zero.