AI Body Shop · jaikey.net

Know what’s going to break, then get it fixed.

Jaikey is the repair shop for prompts, agents, skills (SKILL.md), and workflows. Pull in with what you built, say what it should do, get an honest checkup and a clear ship / fix / rethink call. Nothing gets rebuilt until you say so.

01. What breaks

You’ve shipped broken AI before, or you’re about to.

Solo builders don’t fail for lack of models. They fail because the failure mode only shows up in production, after a user already hit it.

Fear

It “works” in the demo

Happy-path chat looks fine. Multi-turn, bad inputs, or a tool that returns garbage, and the agent quietly does the wrong thing. Users don’t file a bug. They leave.

Cost

Silent failures are expensive

Wrong tool call. Hallucinated handoff. No fallback when the API flakes. You only learn when support tickets, refunds, or a public screenshot show up.

Trap

Chat feels like a review. It isn’t one.

Asking ChatGPT “is this ready?” gets a confident essay, different every time, rarely a hard no. That’s the confidence tax: vibes instead of scores, risks, and a ship / fix / rethink verdict. Frankey is the Evaluator. Chat is not.

02. Projects

Real reviews. Measured before and after.

Not mockups. Full Frankey runs on systems that came out of browser chats and prompt builds. Same report shape. Hard verdict. Score that moves when you actually fix the architecture.

Project: Validation coach prompt

Validation coach prompt

A strong validation agent (Play Matrix, Four Boxes, pre-commit number) that still had a silent-assumption hole and soft REAL RISK. Frankey said repair; after approval, critical and high issues went to zero.

FRK-2026-072114 System prompt, validation coach Confidence 92% → 96%

  • Crit Silent assumptions when inputs missing
  • High REAL RISK free-text → soft-pedaled WTP trap
  • High Play Matrix not coupled to offer levers
  • Med No founder commitment check

What actually broke (and what closed it)

Problem

Silent assumptions

Agent could invent buyer, price, or assets and keep going. Founder never saw what was made up: false “pass” signals. Patch: mandatory INPUT AUDIT as the first output section; every ASSUMPTION → listed with risk.

Problem

Willingness-to-pay soft-pedal

REAL RISK was free text. Agents buried the “free proves engagement, not payment” trap. Patch: rigid Trap / Why a pass would be a lie / Hardening structure, impossible to skip.

Problem

Weak offer after a good play

Play Matrix was excellent; offer sharpening was optional flavor. Patch: VALUE EQUATION LEVERS PULLED must feed the OFFER every run.

Deliverable after approve: Validation_Coach_Prompt_v1.1_OPTIMIZED.md. Full pre/post reports in portfolio/samples/validation-coach-prompt/

Also from browser chats

UI builds that looked shippable in preview

Same builder pattern: something that demos hard in a chat/browser preview, then gets forced through architecture questions: empty success paths, missing product model, trust claims without fail states. Screenshots from internal review loops.

Jobs Raider before: empty resume parse and waiting state Jobs Raider after: active profile and grounded RAG optimization
Jobs Raider: empty parse shell → structured profile + grounded rewrite path
Writer’s Block before: single manuscript with side-panel coach Writer’s Block after: workspace with projects and drafts
Writer’s Block: one-doc coach panel → multi-project workspace
03. Also from Jaikey

Need an Architect - and a private place for it?

That’s Frank - the Architect. Creates and improves prompts and agents into clean build-ready artifacts. Your drafts stay only visible to you, saved in your Frank app. Live as a web app now (Windows console next).

Live · Frank Architect · Private workspace

Paste it. Get a build-ready artifact.

Huge selling point · Your work stays yours Only you can see it. Saved to your Frank app.

Drafts and Architect results are private to your account - not a public feed, not shared with other users, not used to train models for other customers. Build without pasting secrets into a shared chat thread.

Frank is the Architect for ChatGPT, Claude, Gemini, and Grok users who want a clean prompt or agent in seconds - not scores, not findings, not a ship verdict. That’s Frankey’s job as the Evaluator. Paste. Architect. Deploy - from a workspace that’s yours alone.

  • Private by default - only you see your work; saved under your account
  • Create & improve prompts and agents
  • Clean artifact only (no review report)
  • Copy ready-to-use result into any LLM
  • Need a verdict? Open Frankey, the Evaluator
04. Why not just chat

Can you trust an AI to review your AI?

Yes, when review is a fixed engine with a repeatable structure, not another improvisation in a chat window.

General chat is great for creating. It is a weak judge of its own work. Every “review” is a new improvisation with no score history and no hard ship gate. Frankey sits in the gap as the Evaluator - design-time diagnostic review before you ship.

01

Same structure every run, so you can compare

Scores, findings (problem, impact, severity), risks, and a ship / fix / rethink verdict in the same order every time. Re-run after a fix and see what moved, not five different chat essays you can’t reconcile.

02

Judged as the thing it is

A one-line skill is not scored like a multi-agent workflow or a PRD. Frankey classifies the asset and applies the right lens, so “looks fine” is not the default answer for every paste.

03

A verdict you can act on. Nothing auto-patched

Free diagnosis is a second opinion only. Nothing in your system changes unless you explicitly approve a path later. If all you wanted was “can I ship this?”, you can stop at the report.

ChatGPT will happily rate your prompt 9/10. Frankey will tell you what breaks on turn four, which tool call is ambiguous, and whether you should ship, fix, or rethink. Built for the decision “can I trust this?”

How it works

Evaluator pipeline. One job: diagnose.

01 · Load

Prompt, agent, workflow, or PRD

Paste text or load a file. Treated as the system under evaluation - not a chat to continue.

02 · Evaluate

Classify → score → findings

Asset Intelligence, profile-aware scores, risks, and a clear ship / fix / rethink verdict. Frankey is the Evaluator - not a silent rewriter.

03 · Decide

Report first. Fix only if you approve

Diagnosis first. If you accept fixes, Frank (Architect) builds the clean artifact. Nothing is modified without your approval.