myibrahim.cloud

Vibe Coding · Editors & Agents

Vibe coding with Claude Code

How to delegate, when to drive, and what changes about software engineering when an agent is paired with you.

Claude Code is the closest thing today's tooling gives us to a junior engineer who reads the whole repo before touching a line. The shift isn't "AI writes my code". It's a different unit of work. You stop writing functions and start writing prompts that produce, then verify, the function.

That sounds glib. It isn't. The day-to-day actually feels different.

What I lean on it for#

After ~6 months of daily use, here's where Claude Code earns its keep:

  • Mechanical refactors that span a dozen files. Renames, type narrowing, splitting a god-module, migrating from requests to httpx. Things that are tedious-but-mechanical in a way I'd previously have grumbled through with grep + sed.
  • Reading unfamiliar codebases. "What's the auth flow here? Trace it from request to session." The model reads the code and explains it back faster than I would have grepped.
  • First drafts of routine code. REST endpoints, Django serializers, test scaffolds, terraform for a new module. Boilerplate evaporates.
  • Proofreading my own diffs. Paste in a 200-line PR, ask "what would a reviewer flag?" Catches the silly stuff before a human has to.

What I keep manual#

Just as important — the things I don't delegate:

  • Architecture. The model can produce a plausible plan. Choosing between two plans, when both compile, is judgment. That's the thing I get paid for.
  • Anything in production-shaped systems where I haven't already vetted the constraints. I'd rather hand-write a migration than trust an agent to know I have read replicas, a daily snapshot SLA, and a particular ORM quirk.
  • Naming things. Still mine. Names are the first place a future reader looks; they need to come from someone who understands the intent.
  • Reviewing the model's own output. Always. Every time. Especially when it looks right.

A worked example: refactoring a Django view#

Here's a real one. I had a 280-line get_dashboard view that had grown a layer of feature-flag conditionals over a couple of years. I asked Claude Code to "split this into a service layer + a thin view, keep behavior identical, but move the feature-flag branches into a strategy pattern."

Before:

def get_dashboard(request):
    user = request.user
    flags = get_flags(user)
    if flags.get('beta_charts'):
        charts = build_beta_charts(user)
    else:
        charts = build_legacy_charts(user)
    if flags.get('compact_layout'):
        layout = 'compact'
    else:
        layout = 'standard'
    # ...230 more lines like this
    return render(request, 'dashboard.html', {...})

After (model-generated, lightly edited):

# services/dashboard.py
class DashboardBuilder:
    def __init__(self, user, flags):
        self.user = user
        self.flags = flags

    def build(self) -> DashboardContext:
        return DashboardContext(
            charts=self._charts_strategy().build(),
            layout=self._layout_strategy(),
            # ...
        )

    def _charts_strategy(self) -> ChartStrategy:
        if self.flags.beta_charts:
            return BetaChartStrategy(self.user)
        return LegacyChartStrategy(self.user)

# views.py
def get_dashboard(request):
    flags = get_flags(request.user)
    ctx = DashboardBuilder(request.user, flags).build()
    return render(request, 'dashboard.html', asdict(ctx))

The agent's pass got me 80% there. The remaining 20% — naming the strategy classes correctly, deciding which flag groups deserved their own strategy vs staying inline, removing one branch that turned out to be dead code — that's where I earned my hour.

The discipline that actually matters#

If you take one thing from this:

Read the diff. Every time. The output looks reasonable in 95% of cases — and quietly wrong in 5%. The ones that bite you are the ones that look right.

Treat every change like a colleague's PR, not your own. That mindset shift is the entire game.

A few specific reading habits I've built:

  1. Run the tests after every accepted change. Not just at the end — every change. The agent will sometimes confidently rewrite a function whose existing behavior is being relied on by something three modules away. The test suite is your contract enforcement.
  2. Search for the function's callers before accepting a signature change. grep "function_name(" -r src/ is non-negotiable before approving a rename or arg reorder.
  3. Watch for hallucinated APIs. Lower-frequency in 2026 than in 2024, but still happens for niche libraries. If the code calls a method you've never seen, look it up.
  4. Reject "improvements" you didn't ask for. If you asked for a rename and the diff also includes a refactor of error handling, push back. Scope discipline keeps reviews tractable.

The team-level question#

The bigger shift isn't personal — it's organizational. When every engineer on the team is 1.5–3× more productive at mechanical tasks, the bottleneck moves. It's no longer "can we ship the code?" It's "do we know what to build?" and "can review keep up with output?"

I don't have a clean answer for that. The teams I've watched succeed have leaned harder on:

  • Tighter PR review SLAs — review is now the constraint, treat it like one
  • More aggressive deletion — when code is cheap to write, code is cheap to delete; don't let it pile up
  • Better issue hygiene — "the agent will figure it out" doesn't survive contact with vague requirements

If your team is shipping more code but the same number of features, you've reorganized the bottleneck without removing it. That's a sign to look upstream.

What's next#

I'm watching three things in 2026:

  • Long-running agents — Claude can already run for ~30 minutes on a task; sustained multi-hour engineering sessions feel close.
  • Better repository-wide reasoning — "what's the best place to add this feature, given how the codebase is laid out?" is still hit-or-miss.
  • First-party CI integration — running the agent in a worktree, on a branch, with the test suite as the reward signal, is where the real gains are.

For now: pair with it daily, read every diff, and don't let the seductive plausibility of generated code skip your review brain. The 5% wrong is what you're paid to catch. `.trim(), };

  • ai
  • tooling
  • workflow
  • claude
Need this built? I build ai product or mvp projects for clients worldwide. Tell me about yours.