+1 (405) 383-8943

AI coding assistants in a mature Ember codebase

Every team we work with is now running some kind of AI coding assistant — Copilot in the editor, Claude Code or Cursor in a terminal, a review bot on pull requests. On a mature Ember app the results are uneven in a specific, predictable way. The model is fluent in Ember, but it is fluent in the Ember of 2016. It writes Ember.Object.extend, it reaches for this.set, it suggests controllers and mixins and observers, and it does all of it confidently. On an eight-year-old codebase that still contains classic code, nobody notices for a while, because the suggestion looks exactly like the file next to it.

That is the whole problem, and it is worth stating plainly: assistants pull your codebase toward its own average age. Left alone, they will quietly re-add the patterns you are paying to remove. Used deliberately, they are genuinely good at the mechanical half of modernization work. Here is how we set them up on client engagements.

Why the output skews old

Training data is public code and public writing, and the Ember corpus on the public internet peaked years ago. There is vastly more text describing didInsertElement than {{did-insert}} or a modifier, more about ember-cli-mirage than about MSW, more about the Broccoli build than about Embroider or Vite. The newer the Ember idiom, the thinner the evidence: template tag files (.gjs/.gts), the Ember Data request layer, v2 addon format, and Polaris-era APIs are all underrepresented relative to how much they matter in current work.

Two consequences follow. First, the default suggestion is a classic-Ember suggestion. Second — and this is the one that costs real hours — the model will invent plausible APIs in the thin areas. It will confidently produce a @glimmer/component lifecycle hook that does not exist, or Ember Data builder arguments in the wrong shape. The failure mode is not gibberish; it is competent-looking code that does not run.

Give it the project's idioms in writing

The highest-leverage thing you can do takes an afternoon: write down your conventions in the rules file your tool reads (CLAUDE.md, .cursorrules, a Copilot instructions file — the filename changes, the content doesn't). Keep it short and prescriptive, because long rules files get ignored in practice. Ours usually say roughly this:

  • Ember version and LTS line, and that suggestions must target it.
  • Glimmer components with native classes and @tracked; never Ember.Object.extend, this.set, computed properties, mixins, or observers in new code.
  • Modifiers and {{on}} instead of didInsertElement and {{action}}.
  • Services over controllers for shared state; controllers exist only where query params require them.
  • Which test style is current (setupRenderingTest, render, click from @ember/test-helpers) and that legacy test idioms are not to be copied.
  • The addons that are endorsed, and the ones that are being removed — name them, so the model stops suggesting the dying one.
  • Where the template tag is in use, if it is.

Then point the tool at the authoritative sources rather than its memory: the current API docs and the Ember Guides for the version you are on. Most assistants will fetch or index documentation when told to; this is what moves the thin-evidence areas from guessing to reading.

Where it earns its keep

The useful framing is that an assistant is fast at tedious, verifiable work and unreliable at novel, hard-to-check work. On a mature Ember app, three categories of work fit the first description well.

Clearing the deprecation list. You have an inventory from ember-cli-deprecation-workflow and a few hundred call sites. The transformation for each deprecation ID is mechanical and the same every time. Give the assistant one worked example — a merged PR that fixes five instances — and have it do the next thirty under review. When a codemod exists, run the codemod instead; the official codemods are deterministic and the model is not. But for the long tail that no codemod covers, this is where the hours go and where the help is real.

Octane conversion of leaf components. Classic component to Glimmer component is well-trodden ground in the training data, which here works in your favor. Convert one component per PR, with the rendering test written or updated first, and read every diff. Watch specifically for silently dropped two-way binding: classic code that mutated an argument becomes a Glimmer component that no longer does, and nothing fails at build time.

Explaining code nobody owns. This is the underrated one. A twelve-year-old app has routes no current employee has read. Asking for a walkthrough of a route, its model hooks, and everything it touches is a fast way to rebuild context — and unlike generated code, a wrong explanation is cheap, because you verify it against the file in front of you.

Where it does not earn its keep: architecture decisions, anything touching the build pipeline, Embroider configuration, and migration strategy. Those are low-volume, high-consequence, and poorly represented in training data. Do them yourself.

The review bar does not move

Three rules we hold to, and ask clients to hold to.

Nothing merges unreviewed by a human who can explain it. This is not a ceremony. Generated Ember code tends to fail in ways that pass CI: a @tracked field that is set but never read because the template still reads a stale getter, a test that renders nothing and asserts nothing, an ember-concurrency task converted to a plain async method with the cancellation semantics quietly gone.

Lint and type-check before you read. ESLint with the Ember plugins, ember-template-lint, and Glint if you have TypeScript will catch a meaningful fraction of the anachronisms automatically — and unlike a rules file, a lint rule cannot be ignored. If you want one concrete investment here, it is this: every convention you care about that can be expressed as a lint rule should be a lint rule. That is good hygiene for human contributors too, and it is the only mechanism that enforces itself.

Keep the PRs small. The temptation is to let the tool convert forty components in one sitting. Review quality falls off a cliff after the first few hundred lines of diff, and a bad merge in a business-critical app costs far more than the hours saved.

What this does not fix

An assistant does not make an unmaintained app maintainable, and it does not substitute for someone who knows what the app should become. It lowers the cost of the mechanical work — which, on a modernization program, is a large share of the total — while leaving the judgment where it was. If your Ember app is three LTS lines back, the plan still has to be built by a person who has done it before. The assistant just makes the burn-down faster.

If you're partway into a modernization and want a second opinion on sequencing, an assessment is usually the cheapest way to get one.