Principal Harness Engineer
fico · Work from Home, United Kingdom · Remote
Posted Aug 12, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
FICO (NYSE: FICO) is a leading global analytics software company, helping businesses in 100+ countries make better decisions. Join our world-class team today and fulfill your career potential!
The Opportunity
Come join our engineering team in a hands-on technical role at the heart of a new discipline: Harness Engineering. As AI coding agents take on more of the software lifecycle, the hard part is no longer writing code - agents generate it faster than humans can review it, so the bottleneck shifts to verification and trust. Harness Engineering exists to break that bottleneck: engineering the environment that steers agents toward correct, maintainable, well-architected output so that quality is enforced by the system, not re-audited by a person on every change. We call that environment the harness (Agent = Model + Harness). As a Senior Harness Engineer you'll independently own whole harness subsystems, set the standards other engineers build to, and be involved in the end-to-end lifecycle of turning raw model capability into production-grade engineering.
What You’ll Contribute
Design, build, deploy, and support core components of the harness - the guides, feedback loops, guardrails, and shared context that turn raw model capability into production-grade engineering. This is a hands-on role focused on systems and leverage, not hand-writing application code.
Own and evolve feedforward guides - agent instruction files, reusable skills, architectural rules, reference docs, and codemods - and drive team-wide standardisation so agents get it right the first time.
Build feedback sensors - custom linters, static analysis, structural and architecture-fitness tests, verification loops, and LLM-as-judge reviewers - that catch issues automatically before they reach human reviewers.
Own quality gating and release criteria for agent-produced work, defining authority boundaries for what agents may merge unaided and the escalation rules for what must route to a human.
Establish LLM testing infrastructure and evaluation approaches that ensure AI-generated output meets quality and safety thresholds; apply consumer/contract testing (e.g. Pact) where service integration reliability matters.
Run the steering loop - when an agent repeats a mistake, engineer a control so it can't happen again - and treat repository knowledge (docs, specs, context) as the system of record, fighting drift with continuous garbage collection.
Decide where each control runs in the path to production - fast checks pre-commit, more expensive checks post-integration, and continuous sensors that scan for drift outside the change lifecycle - keeping quality as far left as is economical.
Improve observability into agent work and track the measures that matter - cost per merged PR, time-to-merge for agent-assisted PRs, review velocity relative to PR size, defect escape rate, and agent-PR survival rate - using them to decide where to invest next.
Partner with product and…