Time to accepted change
Does faster generation survive review and rework?
Trace a sample from task to merge: cycle time, review wait, revision rounds, and developer effort. Compare similar work where the data allows.
Measure and improve AI-assisted development across delivery speed, code quality, and cost. Start with a focused engineering audit.
For engineering teams using Cursor, Claude Code, Codex, or Copilot.
Understand what changes after the code is generated.
AI adoption is only part of the picture. We examine the work around it: review, rework, architecture, and the cost of getting a change accepted. Then we help improve the workflow.
For established software teams already using AI coding tools. A fixed scope and fee, agreed before we start.
Start with one team, one repository, and a representative sample of recent changes. Review tooling spend, agent setup, and the path from task to accepted code.
Baseline, prioritized findings, and one agreed workflow improvement. Larger implementation is scoped separately.Review available usage exports, recent PRs, CI history, and developer interviews. Identify what can be measured and where the evidence is incomplete.
BASELINE & EVIDENCE GAPSTrace review burden, retries, context use, and model choices. Rank improvements by likely value, effort, and confidence.
PRIORITIZED FINDINGSImplement one agreed change to agent instructions, context, model selection, or review checks. Leave a repeatable evaluation and a follow-up measurement plan.
CONFIGURATION, RUNBOOK & NEXT STEPSFour connected areas. We investigate the parts that matter to your team, using the evidence you actually have.
One engineering workflow.
Four connected perspectives.
Examine where AI shortens a task and where it shifts effort onto reviewers. Look at task boundaries, PR size, revision loops, and the checks needed before a change is accepted.
These are assessment questions, not claimed client results. We establish a baseline before recommending a change.
Does faster generation survive review and rework?
Trace a sample from task to merge: cycle time, review wait, revision rounds, and developer effort. Compare similar work where the data allows.
What does the new code leave behind?
Inspect PR size, duplication, abstraction reuse, churn, and CI failures. Review the code and its context rather than treating line counts as a quality score.
What are you paying for a usable result?
Combine available tool and model spend with retries and estimated review effort. Separate measured costs from estimates and account for missing telemetry.
PR history alone cannot reliably identify AI-written code or prove that AI caused a change in performance. We combine available tool data, code review, and team interviews, and make uncertainty explicit.
You work directly with Raffay Sajjad: the engineer who assesses your workflow and implements the agreed changes.
Production engineering, agentic development, and model economics in the same conversation.
Technical leadership at Trafilea across architecture, delivery, experiments, and production systems.
Finly AI: mobile, backend, infrastructure, and analytics. Idea through a working product.
Product engineeringHands-on practice with coding agents, context management, model selection, and local models. The focus is accepted, maintainable work.
Founder experience behind Tiercel Labs. Not client logos.
Selected writing on engineering productivity, review burden, and the economics of AI-assisted work.
All 28 articles
A controlled Copilot task was 55.8% faster. Field experiments at Microsoft and Accenture show smaller, noisier PR gains.
ReadTotal cost is inference plus retries plus human review plus delay. A cheap model can still be expensive.
ReadToken price is easy to compare. Cost per accepted task is what decides whether a model choice is cheap or expensive.
ReadCTOs, VPs of Engineering, engineering directors, and platform or developer-experience leads at established software teams already using AI coding tools. You have a real workflow to examine and an owner who can help implement changes.
One team, one repository, and a representative sample of recent work. We agree the questions, available evidence, deliverables, timeline, and fixed fee before starting. A focused assessment typically takes one to two weeks after access is ready.
An evidence-backed baseline with its limitations, prioritized findings, one agreed workflow or configuration improvement, and a runbook for measuring what happens next. Larger implementation work is a separate decision.
Usually usage or billing exports, selected PRs and CI history, repository instructions, and conversations with a few developers and reviewers. We agree the minimum access needed. Redacted exports or a guided walkthrough can be used when direct repository access is unsuitable.
No fixed saving is promised before measurement. Repository history alone does not reliably establish AI authorship or causality. We distinguish observed data from estimates and use comparable tasks where possible. Faster generation is not treated as proof of faster delivery.
Yes. The first audit includes one bounded improvement agreed in the scope, such as repository instructions, context configuration, or a review check. Broader tooling and workflow changes are scoped separately after you review the findings.
The first engagement has a fixed fee based on the repository, questions, and access available. We send a written scope and quote before you commit. There is no ongoing retainer required.
No. We start with your existing tools and available data. Cursor, Claude Code, Codex, GitHub Copilot, custom agents, and local models can all be part of the assessment. Tool names describe the workflows we assess, not vendor partnerships.
Tell us which coding tools your team uses and where cost, review, or delivery is becoming difficult to understand.