Metadata-Version: 2.5
Name: pramana-sdk
Version: 0.0.10
Summary: Decide whether a prompt or model change is safe to ship, by running it against your own production runs.
Requires-Python: >=3.12
Requires-Dist: pramana-core>=0.0.2
Requires-Dist: pramana-proto>=0.0.2
Requires-Dist: pramana-store>=0.0.6
Requires-Dist: requests>=2.32
Requires-Dist: zstandard>=0.22
Provides-Extra: replay
Requires-Dist: pramana-store[postgres,s3]>=0.0.6; extra == 'replay'
Description-Content-Type: text/markdown

# Pramana

**Decide whether a prompt or model change is safe to ship, by running it against your own
production runs.**

Pramana is a flight recorder for AI agents. It records every non-deterministic decision your
agent makes in production, replays any past run exactly against your changed code, and reports
which decisions moved — not which sentences got reworded.

## The problem

You change a prompt or upgrade a model and you have no way to know what it broke.

Tests check that your code runs. They never see the decision the model made inside it. Evals
average away a change that breaks eleven cases out of three hundred. And diffing two runs is
useless, because a model never reproduces its own wording — replay one run on a new model and
nearly every step differs, so every diff is red and nobody reads any of them.

## Install

**Requires Python 3.12 or newer.**

```console
$ pip install pramana-sdk
```

On 3.10 or 3.11? Open an issue or email me. It's about half a day of work and nothing architectural blocks it — I'll do it if someone needs it.

That installs what an agent needs to **record**: it posts events to an HTTP sink and writes
payloads to a local directory. Replay and model-diff additionally open a database, so they ask
for an extra:

```console
$ pip install 'pramana-sdk[replay]'
```

The split exists because psycopg and boto3 are ~50 MB that a send-only agent box never executes.

This installs the `pramana` CLI as well:

```console
$ pramana --help
usage: pramana trace show <ref> [--causal-order]
       pramana replay <ref> [<ref>...] -- <command...>
       pramana model-diff <ref> [<ref>...] [--yes] [--change-id=<id>] -- <command...>
       pramana rebaseline <name> -- <command...>
       pramana baseline set <name> <trace_id>
       pramana baseline get <name>
       pramana baseline list
       pramana doctor
```

## Quickstart

Three lines around the agent code you already have:

```python
import uuid

import pramana
from pramana.sinks.http import HttpSink

pramana.init(
    trace_id=str(uuid.uuid4()),
    tenant_id="<your-tenant-id>",
    sink=HttpSink(
        "https://www.reliai.in/ingest/v1/events:batch",
        tenant_id="<your-tenant-id>",
        api_key="<your-api-key>",   # or set PRAMANA_API_KEY and omit this
    ),
)
pramana.instrument(client)          # your existing OpenAI or Anthropic client

# ...run your agent exactly as you already do...

pramana.shutdown()
```

Every call through `client` is now recorded — the model, the prompt, the response, latency and
token counts — without changing how you call it.

## What it records

Model calls, clock reads and random draws are captured automatically once you instrument
your OpenAI or Anthropic client — you do not change how you call the model.

Tool calls are captured where you wrap them with `interceptor.intercepted_call(...)`. That is
a real integration cost and worth knowing up front: the classification is only as complete
as your wrapping, and a tool call Pramana cannot see is a change it cannot report.

```python
from pramana import interceptor
from pramana_proto.v1.event_pb2 import EventKind

result = interceptor.intercepted_call(
    lambda: charge_card(order_id, amount),
    raw_input={"order_id": order_id, "amount": amount},
    kind=EventKind.TOOL_EXEC,
    attrs={"tool": "charge_card"},
)
```

Also captured: state commits and agent-to-agent messages.

## What replay tells you

Every step of a compared run lands in exactly one of eleven outcomes, and the buckets sum to the
step count:

| Outcome | Meaning |
|---|---|
| NO_CHANGE | identical |
| COSMETIC | wording moved, behaviour did not |
| BEHAVIOURAL | the agent did something different |
| CONSEQUENT | diverged only because an earlier step diverged |
| SUPERSEDED | a retry that was replaced by a later attempt |
| REORDERED | same calls, different order |
| HALTED | the run left the recording and stopped there |
| UNSTABLE | ordinary model variance, confirmed by a control run |
| ERRORED | the step raised |
| NO_RECORDING | the call ran, with nothing recorded at that call site to compare it against |
| NOT_CALLED | a recorded call that is no longer in the code |

A call site is identified by the text of the call, so editing the call itself — renaming a
variable inside it — makes it a new site. That shows up as `NO_RECORDING`, not as a changed
decision. A recorded call the run simply *didn't reach*, while the call is still in your code, is
`BEHAVIOURAL`: the agent decided differently, which is the thing worth telling you about.

## Why "cosmetic" is trustworthy here

Cosmetic is decided by elimination, never by reading the text. If every tool call in a run is
identical in name, arguments and order, then nothing the agent *did* changed, however
differently it worded itself.

No embeddings. No similarity score. No threshold to tune. No second model judging the first.
Two engineers looking at the same two runs get the same answer, and so do they six months
later.

A step that diverged only because an earlier step diverged is marked as downstream of the
first, so one root cause is reported once rather than forty times.

## Replaying production runs does not touch your systems

In a sandboxed batch — two or more traces, which is the default — the tool calls you have
wrapped are served from the recording and never executed. The model is called for real,
because the new model is the thing you are testing.

In a plain replay, nothing leaves the process at all: outbound network access is blocked at
the process level, below whatever HTTP client you use.

In both, a tool call whose arguments do not match the recording stops that run and reports
the divergence. Pramana never invents a tool response to keep a run going. A system that
guesses hands you a clean-looking report about a run that never happened, and you cannot
tell which report that was.

One thing to know before you run a batch: Pramana only intercepts tool calls you have
wrapped. An unwrapped call is invisible to it and will execute normally. Wrap the ones with
consequences first.

## Telling a real regression from model noise

A behavioural finding on a model swap is ambiguous: it could be the new model, or the old one
being non-deterministic. So a finding triggers a control run — the same model, replayed again.
If the old model wanders the same way on its own, the finding is labelled unstable instead of
shown to you as a regression.

During a model diff, clock reads and random draws are served from the recording, so the model
is the only thing that changed.

## Evidence

A run exports as a signed bundle. `pramana-verify` checks it offline, with no account and no
network call, and will not accept the public key shipped inside the bundle it is checking.

A valid signature proves nothing has been altered since signing. It does not prove that what
was captured was everything that happened. This is a proof of record, not a judgement of
conduct.

## Limitations

- **You cannot import existing conversation logs.** Replay needs the execution trace, and a
  transcript does not carry one. Your corpus starts the day you instrument.
- **It reports that a decision changed, never whether the change is good.** Judging a changed
  decision is your call.
- Python only. No JavaScript or TypeScript SDK.
- The OpenAI and Anthropic clients are supported. Azure OpenAI and `AnthropicBedrock` are routed
  to those adapters and covered by tests; neither has been exercised against live cloud
  credentials. Raw `boto3` is not supported.
- LangChain is tested. LangGraph, CrewAI and LlamaIndex are not, and are not supported.
- No SOC 2, no penetration test, no uptime SLA, no on-premise deployment.

## If what you want is tracing

Use Langfuse. It is open source, free, and better at tracing than we will be. Pramana answers
a different question: what did your change do.

---

[Documentation](https://www.reliai.in/docs/) · [reliai.in](https://www.reliai.in/) ·
[`pramana-verify`](https://pypi.org/project/pramana-verify/)
