🌱 Beginner · Guide 1 of 2
What Vibe Coding Can and Cannot Do
9 min read · Last reviewed 23 Sep 2026

You have heard that you can now describe an app and get one. That is partly true. The part that is not true is where people lose weekends, lose money, and occasionally lose other people's data. This guide is the capability map that nobody selling a subscription has much reason to publish: what these tools genuinely do well, what they do badly, the failure modes you will actually meet in your first month, and where the honest ceiling sits.
Nothing here assumes you can program. If a word is unfamiliar, the glossary has a plain definition.
What they reliably do well
First drafts of small, self-contained things. A landing page. A form that emails you. A script that renames three hundred photos. A small web app with a handful of screens and no accounts. These are jobs where a good answer looks much like a million other good answers already written down, and where you can see within seconds whether it worked.
Boilerplate and glue. Configuration files, the plumbing that connects one library to another, the tedious parts that experienced programmers also find tedious. A model that has read a great deal of code is very good at the parts of code that are nearly identical everywhere.
Explaining code you did not write. This is the single most underrated use for a beginner. Paste a file in, ask what it does, ask what breaks if you change line 40. You get a patient explanation at whatever level you ask for, and you can check it against reality by changing line 40.
Mechanical edits across many files. Rename a thing everywhere. Change a date format in twelve places. Boring, error-prone work that agents do faster than you and roughly as accurately.
Writing and running tests. A test is a small piece of code that checks another piece of code still does what it should. Agents write them quickly, run them, read the failures and fix them, and that loop is the main reason an agent can work unsupervised for more than a minute at a time. One honest caveat: if the agent misunderstood what you wanted, it will write tests that confirm its misunderstanding and report success. Tests catch the code breaking later; they do not catch the goal being wrong now. That part is your job, and it is the job you are actually qualified for.
Getting unstuck. Paste an error message you do not understand. Even a wrong answer usually contains the vocabulary you needed to search for.
Claude Code Paid
Anthropic's agentic coding tool for the terminal, IDE, desktop and web
🔥 Claude Pro billed annually: $17/month instead of $20Where they are unreliable
Architecture that has to survive. An agent optimises for the next reply, not for month six. It will happily give you a structure that works today and becomes impossible to change later, because "impossible to change later" is not visible in the output you are reviewing now.
Genuinely novel problems. If your problem has no close relative in public code, quality falls off sharply. You will not notice the difference in tone, only in results.
Anything requiring true correctness. Money arithmetic, tax, dates across time zones, anything medical or safety-related, anything where two things happen at once. These tools produce code that is usually right. "Usually right" is a different product from "provably right", and no amount of prompting closes that gap.
Secrets. Left alone, an agent will put a password or an API key directly into a file, and that file will end up somewhere public. It is not being careless on purpose; it has no reliable sense of which strings are dangerous.
Keeping a big codebase coherent. An agent sees only what fits in its working memory. Past a certain size it stops finding the function you already have and writes a second one that does almost the same thing. A few months of that and nobody, human or machine, can safely change anything.
This is not a critic's view. GitHub says the same thing about its own product in its responsible-use documentation: Copilot Chat can produce code that looks valid but is not, answers that sound plausible and are wrong, and code that exposes sensitive information or security holes if you do not review it. Every serious vendor publishes some version of that warning, usually on a page you have to go looking for.
Why the good list and the bad list look the way they do
There is one rule underneath both lists, and it is worth internalising because it predicts new tools you have not tried yet. These systems are excellent when the work is common, small and checkable, and unreliable when it is rare, large or only checkable much later.
Common, because the pattern has to have existed somewhere in what the model learned. Small, because everything the agent is currently considering has to fit in a working memory that is generous but finite. Checkable, because the agent improves by seeing a result — a test failing, an error message, a page rendering — and where there is no fast feedback there is nothing to correct against.
Score your idea against those three words before you start and you will predict your own experience surprisingly well. "A page that shows my Instagram photos in a grid" scores three out of three. "A scheduling system that must never double-book across time zones" scores zero, and the failure will arrive in production, quietly, six weeks later.
The five failure modes you will actually meet
1. Confident wrong answers. There is no tone difference between a correct answer and an invented one. The tool cannot flag its own uncertainty reliably, so you cannot use confidence as a signal at all. Assume every factual claim about a library, a price or an API is a guess until something checks it.
2. Silent breakage. You ask for a change on page three. It arrives, it works, and page one quietly stopped working. You find out four changes later, by which point you cannot tell which change did it. The fix is dull and non-negotiable: put the project in version control on day one, and commit every time something works. "It worked an hour ago" is worthless; "it worked at this exact saved point" gets you back.
3. Drift. Twenty messages in, you are building something adjacent to what you set out to build. Each step was reasonable; the sum was not. Writing the goal down in a file the agent reads every session, and asking for a plan before any code, both help. So do skills libraries that force the agent to design, plan and test in order rather than diving in.
Superpowers 🧩 SkillFree
A skills library that makes coding agents plan, test-first and debug systematically
4. Dependency and security sludge. Agents reach for a library to solve problems you could solve in six lines, and each one you accept is something you now have to keep updated forever. They also reproduce insecure patterns from their training data. Lovable, a builder aimed squarely at non-programmers, says it plainly in its own documentation: misconfigured row-level security rules are "a common cause of data leaks", and its built-in scans "do not replace a thorough security review".
5. Cost surprises. Two shapes. On builders, you spend a daily or monthly token allowance far faster than you expect, because a failed attempt costs the same as a successful one. On terminal agents, a session left open all day keeps getting more expensive, because the whole conversation is re-sent with every message. Anthropic's own cost documentation says as much, and puts the average across enterprise deployments at "around $13 per developer per active day and $150-250 per developer per month". You are not an enterprise developer, but the shape of the curve is the same. Our pricing pages and the monthly cost breakdown on the blog cover what people actually pay. A local usage dashboard is worth installing early, before the first surprising invoice rather than after.
Codeburn Open source
Free local dashboard that tracks AI coding token spend across 40+ tools
The honest ceiling: when "it works on my screen" stops being enough
For a personal tool, a prototype, a demo, a site with no login and nothing to lose, "it works on my screen" is a complete and sufficient standard. Ship it. Enjoy it.
There are four doors, and behind each of them that standard fails:
Data you would be upset to lose. The moment something stores information you cannot recreate, you need backups you have actually restored from once, and a plan for changing the shape of stored data without destroying it. Agents are poor at this because the damage is invisible until it is permanent.
Accounts and logins. A broken login page is visible. A login that works but lets any signed-in person read any other person's records looks identical from your screen. This is the single most common serious defect in apps built this way, and you will not find it by clicking around your own app.
Payments. Never hand-roll this. Use a payment provider's hosted checkout so that card details never touch code you or an agent wrote.
Other people's personal data. Once real users' names, emails, messages or photos are involved, you have legal obligations that vary by country and are not a coding problem. This guide is not legal advice; it is a flag that the question exists before you launch, not after.
The builders aimed at non-programmers know this, which is why they bundle a database, a login system and security scanners rather than leaving you to assemble them. That genuinely lowers the odds of the worst mistakes. It does not move the ceiling, because the responsibility stays with you: Lovable's documentation is explicit that "You are responsible for ensuring that your app meets the security requirements appropriate for its use case."
Lovable Freemium
Build full-stack web apps by chatting, with Supabase and one-click deploy
🔥 Students & teachers: 50% off Pro for up to 12 monthsThere is a fifth thing that applies everywhere: an agent that can run commands and read web pages can be steered by whatever it reads. Anthropic's security documentation describes its safeguards and then states the honest conclusion: "no system is completely immune to all attacks". The same page notes that Anthropic "does not security-audit or manage any MCP server", which is worth remembering every time you add one from a stranger's repository. Our ecosystem map shows what connects to what; it does not vouch for anyone's code, and neither should you.
How to work inside the limits
None of this means "do not use these tools". It means use them like power tools rather than like a colleague.
- Version control from message one. Not later. Every builder and every agent supports it, usually by connecting a GitHub account.
- Ask for the plan first. Read it. It is in English, you can judge it, and it costs nothing to reject.
- Make it prove the thing works. Ask for the exact steps to check, then do them yourself. Ask it to run the tests and show you the output, not to report that tests passed.
- Start fresh sessions often. Long sessions drift and cost more. When the subject changes, start again.
- Cap the money before you start, using whatever spending limit your tool or account offers.
- Treat accounts, payments and other people's data as a stop sign, not a speed bump. Use a hosted provider for each, or get a person who does this professionally to look before you launch.
What to do next
- Pick one small, self-contained, low-stakes project. Something you would be mildly pleased to own and untroubled to lose.
- Choose a tool in ten minutes rather than ten days: Choose Your First AI Coding App walks through the four questions that decide it.
- Put the project in version control before you ask for the second feature.
- When you have shipped that first thing and want a repeatable way of working, move up to the solo builder guides. The ideas that carry across tools — instruction files, skills, MCP servers — are laid out in what each one does.
Read the official docs for…
- What a vendor admits its own tool gets wrong — GitHub's responsible-use page for Copilot Chat, which lists inaccuracy, hallucination, insecure output and bias in GitHub's own words.
- Permissions, sandboxing and prompt injection — Claude Code security.
- Why a long session gets expensive, and how to cap spend — Claude Code cost management.
- What a no-terminal builder does and does not check for you — Lovable's security documentation.
- How your code is stored and whether it trains a model — Cursor's security page.
Mentioned in this guide
Claude Code Paid
Anthropic's agentic coding tool for the terminal, IDE, desktop and web
🔥 Claude Pro billed annually: $17/month instead of $20Superpowers 🧩 SkillFree
A skills library that makes coding agents plan, test-first and debug systematically
Lovable Freemium
Build full-stack web apps by chatting, with Supabase and one-click deploy
🔥 Students & teachers: 50% off Pro for up to 12 monthsCodeburn Open source
Free local dashboard that tracks AI coding token spend across 40+ tools