How we work

The method is the product.

Every engagement follows the same shape, because the shape is what prevents the failures the research describes: pilots that never reach daily use, graphs that stall at entity resolution, and data that decays as soon as it is loaded.

Seven steps

Checked before the next begins.

In order, every time.

01

Count first

Before any processing, we inventory what you have: how many files, of what type, how many pages, how many are duplicates by content. We confirm the numbers with you. A surprising share of “we have 500 documents” turns out to be 380 unique ones and 120 copies, and the estimate, the price, and the plan change accordingly.

02

Sample before scale

We process a representative sample and show you the result: extraction quality, the records produced, the gaps. If your documents come in six formats, we say so and adjust before the full run rather than discover it at the end.

03

Define the structure with you

The ontology (sites, assets, contracts, whatever your domain needs) is agreed in writing before extraction starts. Field lists are explicit. Every field is either filled from the document or left blank; nothing is guessed to make a table look complete.

04

Dry run before any write

No record is written to your systems until you have seen exactly what will be written: how many, to which parents, which fields, and which rows could not be matched and why.

05

Verify from the target

After a write, we query the target system and report its counts, not ours.

06

Report the failures

Every run produces a report of what could not be processed, what was flagged for review, and what was skipped. It is in the summary, not in an appendix.

07

Leave it runnable

Everything is a repeatable, documented pipeline. Nothing is a one-off.

What we do not do

Four things we will decline, in writing.

Saying no early is cheaper than saying sorry late. These are the lines we hold on every engagement.

We do not train models on your data

Or anyone else’s. We use commercially available models under contract terms that exclude training on your inputs, and we choose models per task on accuracy and cost, not brand.

We do not scrape behind logins

Or bypass rate limits or bot protection, or collect personal data without a lawful basis. We will decline the work.

We do not sell dashboards of unverified data

If we cannot count it, we will not chart it.

We do not deliver a demo and leave

Every engagement ends with something running in your environment, documented, with your team able to operate it.

Engagement models

Start small. Each step useful on its own.

Fixed prices where the scope is known. Real numbers before the scope grows.

Assessment

Start here

1 to 2 weeks · Fixed price

We inventory your sources, sample them, and return a written report: what is there, what can be extracted with what confidence, which questions it could answer, what it would take. Useful on its own even if nothing follows.

Pilot

3 to 6 weeks · Fixed price

A bounded slice: a few hundred documents, one or two API sources, one target system. You get the extraction quality report, a working assistant over the sample, a dry run against a sandbox of your target system, and a proposal for the full build with real numbers.

Build

Delivered in stages · Quoted after the pilot

The full scope, each stage accepted before the next starts. Priced on volume (documents, pages, records, sources) and complexity.

Operate

Monthly · Includes a fixed allowance of refinements

New sources processed as they arrive, assistants kept current, connectors maintained through upstream changes, monitoring, and a monthly summary of what changed.

Advisory

As needed · For teams building it themselves

Ontology design, pipeline review, evaluation design, and an honest opinion on what will and will not work.

Count first. Then decide.

Book an assessment. In one to two weeks you will know what you have, what it can answer, and what it would take. Fixed price, and yours to keep whatever you decide next.