Get the data out.
Get it right. Get it connected.
Keep it that way.
Everything we deliver falls into one of five layers. Clients buy one layer or all five. Each is designed to be useful on its own and better with the others.
Acquire
1,284 files · 6 formats
Count
380 unique · 120 copies by content
Sample
40 documents · quality report issued
Structure
7,412 records · 3 entity types
Dry run
7,388 to write · 24 unmatched, listed
Acquisition
Getting the information in.
We read from where the data already lives. We do not ask you to email us copies.
Documents
PDFs, scans, photographs of paper, spreadsheets, presentations, email archives, CAD exports. Converted to text, tables and images with layout preserved: a table stays a table, a caption stays with its photo, a multi-column page reads in the right order. Scanned and low-quality originals are the normal case for us. Every page is accounted for and every unreadable one is reported rather than silently dropped.
APIs
CRM, ERP, ticketing, finance, project management, building management, IoT platforms, and any vendor system with a programmatic interface. We build the connectors, handle authentication, pagination, rate limits and schema changes, and land the data in a form that can be joined to everything else. Each data flow is scoped separately, delivered working, and only then is the next one added.
Web and public data
Supplier catalogues, regulatory registers, competitor pricing, tender portals, public filings, product specifications, news. Collected on a schedule, change-detected, structured. Within the law and within the site’s own rules: public data, no bypassing of technical barriers, terms of service respected, personal data handled under the applicable privacy regime. Where a site offers an API or a licensed feed, we use that instead.
Databases and files
Legacy databases, exports from retired systems, shared drives, cloud storage. Read in place.
Structure
Making the information mean something.
This is the step most projects fail at, and it is where we spend the most care.
Extraction to records
Free text becomes typed records. A purchase order becomes line items with quantity, part number, price and ship date. A contract becomes parties, dates, obligations and terms. An inspection report becomes findings with locations and severities. The structure is defined with you before extraction starts, and it follows your business, not a generic template.
Ontology and knowledge graph
Records are connected: which order bought which item, which item is installed at which site, which warranty covers it, which supplier made it, which person signed off, which document says so. The vocabulary is written down as an ontology you own. Graph-backed retrieval wins on questions that span several sources; plain search wins on single-fact lookups. We build both and route each question to the right one.
Entity resolution and de-duplication
The same site under four names. The same product with and without a hyphen. The same purchase order twice because the package contains both the order and its acknowledgment. We find these by content, not by filename, merge what should be merged, keep what should stay apart, and record every merge so it can be undone.
Reconciliation across systems
Your CRM, finance system, project files and field tools each identify a site or customer differently. We build the crosswalk, show you where the systems disagree, fix what can be fixed automatically, and hand you a short list of the rest for a human decision.
Access
Letting people and systems use it.
The point of clean data is to put it where people already work.
Search with citations
Every page of every document is indexed for meaning, not just keywords. A question about a product family finds the specification sheet even when the model number is misspelled. Every result carries the document name and page. A person can verify anything in one click, which is the only reason people come to trust a system like this.
Assistants and agents
Built on your data only, for specific jobs: a service desk assistant that knows what is installed where; a procurement assistant that knows contract terms and supplier history; an intake assistant that reads incoming documents and files them correctly. Each has a fixed set of tools, a fixed set of data sources, and written rules about what it must never disclose, enforced in the system, not left to the model’s discretion.
Writing back to your systems
Records written into your CRM, ERP, asset register or maintenance system as proper objects under the right parent, with a stable key so re-runs update rather than duplicate. Every write is preceded by a dry run you can read, and followed by a verification query against the target so the count you see comes from the system itself.
Exposing your data to your own AI tools
Your graph, your search and your workflows published as tools any compliant assistant can call, with your access controls in front, using the open standard now supported by every major AI vendor. We build and host them.
Dashboards and exports
Sometimes the right output is a spreadsheet or a dashboard. We deliver those too, from the same governed data, so the numbers match everywhere.
Automation
Making it happen without a person.
Visible, editable, monitored. Not scripts on someone’s laptop.
Workflows
Document arrives, gets extracted, structured, checked, written to the CRM, and someone gets a summary. Or a supplier’s public price list changes, the difference is computed, and procurement gets a note. Built as visible, editable workflows on automation platforms your team can open and understand.
Agents with guardrails
Where a task needs judgment, an agent does the judgment and a workflow does the rest. Agents have hard limits on what they can touch, log every tool call with who asked and what came back, and stop for a human when the action is irreversible or the confidence is low. Logged decisions, human oversight and explainability are what auditors and regulators expect. We build it in from the start.
Scheduling and monitoring
Everything runs on a schedule, with alerts when a source changes shape, a connector fails, or a run produces numbers outside the expected range.
Operation
Keeping it alive.
If you stop working with us, you keep everything and can run it yourself.
Hosting
Pipelines, graph, search indexes and assistants run on infrastructure we manage, writing to storage and databases in accounts you own. Compute-heavy stages are scaled to your volume and switched off when idle.
Re-runs and refreshes
Documents keep arriving, sources keep changing, ontologies get refined. Every stage is repeatable and persists its output, so a refinement at stage four does not mean re-paying for stages one to three.
Handover
Every deliverable comes with documentation a competent engineer can follow, and every dataset is exportable in open formats: plain text, JSON, CSV, a graph export.
For the person who will be asked whether this is sound.
Document parsing
Layout-aware models emit text, tables and images with positions, followed by domain-specific cleanup and cross-referencing. Outputs are persisted per stage so any stage can be re-run alone.
Chunking
Structural, not fixed-length. Table rows, sections and captioned figures become units, each carrying breadcrumbs (document, section, page, entity) so retrieval returns context, not fragments.
Retrieval
Dense vector search, keyword filters on typed metadata, and graph traversal. Questions are routed to the method that answers them best. Results carry provenance.
Graph
Property-graph or RDF storage depending on your ecosystem, with an explicit ontology, stable URIs per entity, and merge provenance recorded on every node.
Assistants
Tool-using agents with a fixed tool set, per-tool descriptions written for the model, hard-coded confidentiality rules, session memory, and full tool-call logging. Model choice is per task and swappable.
Integration
The target platform’s own authentication and credential store. No secrets in code, files or workflow definitions. Writes use upsert semantics on an external key. Every write path has a dry-run mode and a verification query.
Automation
Workflow platforms with visible, versioned definitions, retries and alerting.
Data residency
In the client’s accounts and region. Export in open formats on request.
Compliance
Designed in: logging and human oversight for agents, disclosure where AI-generated content reaches the public, privacy handling for personal data, and a written record of data sources and their terms.
See it on your own documents.
An assessment inventories your sources, processes a sample, and tells you in writing what can be extracted with what confidence. One to two weeks, fixed price.