Research product. We are the operator.

Adjoint is a small AI lab. We build language-model systems — agents, retrieval, and the evaluation infrastructure that keeps them honest — and put them into production alongside a few research labs and ambitious startups.

Founded2024
Team3 engineers & researchers
EngagementsSelective · 3–4 active
StatusSelectively taking new work
§ 02 — Practice

Three threads, woven together. Each pulls the others tight.

Most teams hit a wall when a prototype meets reality. Latency budgets, regressions on real users, evaluation that doesn't generalize. Our practice is built around the joints where those problems live.

01 / Agents

Agents that ship

Tool-using systems that survive contact with messy production data. We design the planner, the memory, the failure modes — and the off-ramp when the model is wrong.

  • Planner & tool-routing architectures
  • Long-horizon task decomposition
  • Human-in-the-loop interrupts
  • Cost & latency budgets
02 / Retrieval

Retrieval that holds up

Hybrid retrieval pipelines that actually answer the question. We treat indexing, ranking, and reranking as a single system — measured end-to-end, not piece by piece.

  • Dense / sparse hybrid indexes
  • Domain-aware chunking & OCR
  • Reranker training & distillation
  • Latency-tier serving
03 / Evaluation

Evaluation that keeps you honest

The unglamorous core of every shippable system. We build evaluation harnesses, golden sets, and continuous regression tests — and we'll teach your team to maintain them.

  • Task-specific eval design
  • Golden set curation & review tooling
  • LLM-as-judge calibration
  • Production telemetry → offline evals
§ 03 — Engagements

Selected work, 2024 — 2026.

A sample of engagements, lightly anonymized. We work in long arcs — typical engagements are 3–9 months with embedded engineering.

2026
Eval harness for a frontier-model research team w/ a research lab · ongoing
Built a continuous eval system spanning 40+ task suites, with selective re-annotation and drift detection. Replaced a quarterly manual review with a daily signal.
EvaluationInfra
2025
Retrieval & reranking for a Finance AI startup w/ Seed startup · 8 months
Replaced a vanilla embedding index with a domain-tuned hybrid retriever and a distilled reranker. Cut answer-not-supported errors by 4.3× on the internal benchmark.
RetrievalDistillation
2025
Customer-support agent, end-to-end w/ FinTech startup · 6 months
Designed the planner, tool layer, and human-handoff protocol. Now resolves 41% of inbound tickets without escalation, with a calibrated abstention rate.
AgentsEvaluation
2024
Code-review agent for an open-source maintainer w/ community collab · 4 months
A focused agent that proposes review comments on PRs, calibrated to a per-repo style. Quietly maintained as an internal tool by the maintainer's team.
Agents
2024
Document understanding for an enterprise pilot w/ Series A startup · 5 months
OCR + structured extraction + retrieval over 2M scanned engineering documents. Shipped as an internal product; handed off with eval harness and runbooks.
RetrievalOCR
§ 04 — Who we work with

A few labs. A few startups. That's the whole roster.

We stay small on purpose. Engagements are selective and long-arc; we're more useful as a fractional team than as a vendor.

Research lab
Texas A&M University
Research and teaching assistant agent.
Seed
Torange
Portfolio, expenses, and retirement modeling on a self-maintaining knowledge base.
FinTech
Mingxi Capital
Research & analyst tooling for an investment team.
Open source
Selected maintainers
Code-review tooling, contributed quietly.
§ 05 — Writing

Notes from the production floor.

Short, technical, written for engineers and researchers who have to ship something on Monday.

§ 06 — Contact

Have a problem at the joint of research and product?

We take on a small number of engagements each quarter. A short note about your problem is the right way to start — we'll reply within a week, even if it's a no.