AI-assisted engineering.
Under examination.
Notes on what happens after code generation: review, rework, delivery, and the economics of accepted work.
Engineering effectiveness
Start here: delivery, review effort, model economics, and agent workflows.
Agent demos hide the hard parts.
Tools, permissions, loops, observability, and stop conditions are the product. Demos skip them.
5 min readLocal models when data cannot leave.
On-prem and local inference make sense under hard data constraints. Capability and ops cost are the tradeoff.
5 min readThe cost of retries and review.
Total cost is inference plus retries plus human review plus delay. A cheap model can still be expensive.
5 min readShadow AI needs a lighter governance model.
Approved tools, data classes, review for external actions, and logging beat a policy nobody follows.
5 min readLab speed is not field pull requests.
A controlled Copilot task was 55.8% faster. Field experiments at Microsoft and Accenture show smaller, noisier PR gains.
6 min readMeasure the work. Not the tokens.
Token price is easy to compare. Cost per accepted task is what decides whether a model choice is cheap or expensive.
4 min readResearch archive
Earlier writing on AI products, business workflows, and published industry cases. These articles remain available in their original context.
Where Jev fits in business automation.
Ticket routing, refund gates, alert triage, agent tool choice — high-volume judgments with clear options. Skip it when you still need a paragraph.
5 min readDecision models vs chat models in a real workflow.
Use a general LLM when you need language. Use a decision model when you need a branch. Mixing them up is how bills and latency explode.
5 min readWhat Jev is — and why it spread so fast.
TypeSafe’s System One model returns typed decisions, not chat. Vercel reported the fastest first-day AI Gateway adoption in its history.
6 min readTwo years of AI ROI: what moved.
Across McKinsey, BCG, and Stanford’s AI Index: adoption rose, enterprise EBIT stayed rare, workflow redesign and KPIs mattered.
7 min readLearn It Live cut tickets 40%.
Public Zapier case study: an e-learning team built a chatbot in about an hour and cut support tickets by 40%. Not Tiercel work.
5 min readWriting tasks got faster—and better.
Noy & Zhang (Science): 453 professionals, ChatGPT on writing tasks—about 40% less time and 18% higher quality.
5 min readSprint-based AI delivery.
Pick one workflow, set acceptance criteria, ship a handover, then decide. Fixed scope beats open-ended transformation.
5 min read88% use AI. Few have scaled it.
McKinsey’s November 2025 survey: adoption is common, enterprise EBIT impact is not, and agents are rarely scaled.
6 min readOnly 5% get AI value at scale.
BCG’s 2025 Widening AI Value Gap: n>1,250 companies, a thin top tier, and a large stuck middle.
4 min readRAG fails when retrieval is wrong.
Evaluate the retriever first: citations, golden questions, permissions, and stale documents.
4 min readHackerOne’s scheduling automation ROI.
Public Calendly story: 169% ROI, 588 hours saved, more meetings and customers—with SSO/SCIM constraints.
4 min readKeep a person in the loop.
HITL patterns that actually ship: draft vs send, prepare vs execute, confidence with sources, clear escalation.
3 min readRebrandly cut support tickets in half.
A public Zapier story: docs-grounded chat, Slack+Jira escalation, and 16,000+ resolved conversations.
4 min readAI Index 2025: what businesses should notice.
Adoption and investment jumped. Reported functional cost savings often stayed under 10%.
4 min readWhen not to use AI.
Four plain conditions where a model adds cost, risk, or noise instead of leverage.
3 min readInference costs fell more than 280×.
Stanford’s AI Index 2025 puts a hard number on how fast GPT-3.5-level performance got cheap.
4 min readLitmus consolidated support. Retention rose.
Public Help Scout case study: Litmus moved off split Intercom + WordPress tooling; early support interaction tied to 26% higher retention.
5 min readOne workflow. Then a platform.
A shared AI platform designed before a real workflow exists usually standardizes the wrong things.
4 min readGen AI use doubled. EBIT impact stayed rare.
McKinsey: regular gen AI use hit 65% in early 2024. Later survey work still finds most firms report no tangible enterprise EBIT from it.
5 min readThe prototype is not the product.
A convincing demo is not a system people can depend on. Product work starts when inputs stop being carefully chosen.
4 min readSupport assistants raised resolution 14%.
NBER evidence from 5,179 agents: generative AI guidance lifted issues resolved per hour—especially for novices.
5 min readMost AI pilots stall before value.
BCG finds only 4% of companies create substantial AI value. Nearly half are still stuck in proofs of concept.
5 min read