# Prefactor

> Prefactor scores every agent run for quality, drift and risk in production, then acts on it.

**Website:** https://prefactor.tech/

## Overview

Most agents pass their evals and fail in production. Prefactor is the evaluation layer that closes the gap. We score every agent run in real time, surface quality regressions and drift as they happen, and show engineering teams exactly how their agents are performing at scale. Built for the teams shipping agents to customers. Go to https://app.prefactorai.com/ to our free dev tier with 25k steps now.

## Features

- Real-Time Agent Evaluation
- Observe-Evaluate-Act Loop
- Runtime Risk Enforcement
- Human-In-The-Loop Approval
- Instant SDK Integrations

## Pricing

subscription

## Alternatives

- [Arize AI](https://nextbigproduct.com/alternativeto/arize-ai)
- [Helicone](https://helicone.ai/)
- [Weights & Biases Prompts](https://wandb.ai/site/prompts)
- [LangSmith](https://smith.langchain.com)

## Details

What Prefactor is


Prefactor is an evaluation and monitoring layer for teams running AI agents in production. Agents that pass offline evals often behave differently once they meet real traffic: a model provider ships an update, an input falls outside the test set, a prompt change shifts behaviour in a way the test suite never covered. Prefactor scores each agent run as it happens in production, so you see how your agents behave on live traffic and not only on the cases you thought to test. You instrument your agent with the SDK, and Prefactor records each step of a run as a span, scores the run for quality and risk, checks it against the activity schemas you define, and keeps an audit trail of what the agent did and why each score came out the way it did.


Main features


Real-time scoring: every run gets a quality and risk score as it executes in production, not only in a pre-launch test set.


Regression detection: when a model update, prompt change, or deploy moves scores down, you see which runs and workflows dropped and by how much.


Drift detection: gradual shifts in agent behaviour surface over days and weeks, before they turn into support tickets.


Schema validation: Prefactor checks each run against the activity schemas you define, so a run that strays from the allowed steps is flagged rather than passing silently.


Trace-level inspection: open any run and read it step by step, from input to tool calls to output, to find where a failure started.


Audit trail: every run, its spans, and its scores are retained, so you can reconstruct how an agent behaved on a specific day and explain it later.


Alerts: you set the runs, cohorts, or workflows you care about, and Prefactor tells you when their scores drop, so you hear it before a customer does.


Who can use it


Prefactor is for the people responsible for an agent once it is in front of customers: AI product engineers, applied machine learning and LLM engineers, and platform teams. It fits a team running one customer-facing agent as well as a team running many agents across two or more products. It sits next to your offline evals and CI rather than replacing them: offline evals gate a release, and Prefactor watches what that release does on real traffic.


What it solves


&nbsp;


Offline evals tell you an agent passed a fixed set of cases. They do not tell you how it behaves on the inputs you did not anticipate, which is where production failures tend to appear. Prefactor closes that gap. It turns manual spot checks and anecdotes into a score on every run, so a change that lowers quality shows up in minutes instead of in a churn report weeks later. Because every run is scored the same way, you can compare one agent version against the previous one and tell whether a change helped or hurt before it reaches all of your users. It catches the failures that only appear in production, the rare inputs and the slow drift, and it gives you a record you can point to when someone asks how an agent behaved.