Practice 02 / AI engineering

AI that holds up in production.

Language model systems built the way regulated software is built. Measured accuracy, traceable output, inference in the region you require, and spend that cannot run away.

Back to all services
Why teams call us

The problems worth using AI on.

We are not interested in adding a chat box to a product. The work that pays for itself removes a bottleneck that people can point at.

  1. 01

    Skilled people are re-typing documents

    Reports, forms and statements arrive in a dozen layouts and leave as keystrokes. The cost is not only the hours. It is the errors a compliance record cannot carry.

  2. 02

    The pilot will not survive production

    A demonstration that reads well proves very little. Production needs evaluation against ground truth, failure handling, cost ceilings and an owner for the day it behaves oddly.

  3. 03

    Nobody can check the answer

    An output with no provenance cannot be reviewed, audited or defended. Traceability back to the source is usually what the business trusts first, ahead of accuracy.

  4. 04

    The data cannot leave the country

    Sovereignty, privacy and contractual obligations decide where inference runs. That is an architecture decision, and it has to be made before the build, not after.

  5. 05

    Spend has no ceiling

    Token cost scales with use, and use grows quietly. Caps, rate limits and caching belong in the design, with a worst case that has been quantified rather than assumed.

Capability

What we build,
in concrete terms.

Six areas of work. Most engagements combine three of them, because a useful system needs extraction, retrieval and evaluation together.

LLM applications

User-facing systems built on language models, with the surrounding product engineering that makes them usable rather than impressive.

Prompt architectureStructured outputReview interfacesStreamingHuman in the loop

Document intelligence

Turning inconsistent documents into structured records. Layout-aware parsing, per-format profiles and validation against the target schema.

Table extractionSchema inferenceFormat profilesCell-level provenanceValidation rules

Agents and automation

Multi-step systems that call tools and take actions, bounded by explicit permissions and an audit trail of what they did.

Orchestration graphsTool useState and memoryGuardrailsEscalation paths

Retrieval and knowledge

Answers grounded in your own material, with the citation attached, so a reviewer can go straight to the source.

Chunking strategyVector and hybrid searchRerankingCitationAccess-aware retrieval

Evaluation and assurance

Measured accuracy against ground truth, not impressions. Evaluation sets are built before tuning starts and re-run on every change.

Ground truth setsRegression evaluationConsensus enginesError analysisAcceptance thresholds

Responsible, in-region rollout

Inference in a required data region, spend guardrails enforced in policy, and a staged path from pilot to supported service.

Data residencyHard spend capsRate limitingKill switchObservabilityAccess control
Evidence

Document intelligence on a compliance register.

Our extraction platform reads consultant inspection reports and writes structured, source-traceable records into an Australian state government agency's Salesforce org. It is in user acceptance testing.

Document intelligence

What the build had to get right

  • Five different consultant report layouts run through one configurable extraction pipeline. New providers are onboarded by configuration.
  • Every extracted value traces back to its page and bounding box in the source PDF. The compliance team trusted the bounding box before it trusted the model.
  • Nothing reaches Salesforce until a reviewer confirms it, and output is validated against picklists, dependency rules and record validation before export.
  • Inference runs on Australia-resident model profiles in an Australian workload region, with observability onshore.
  • Spend guardrails were designed before the first production extraction: IAM-enforced hard caps, per-model rate limiters, prompt caching and a kill switch.

An optimisation campaign on a 190-page benchmark survey cut end-to-end processing from about three hours to roughly half an hour, and per-document inference cost by close to ninety per cent. These are controlled build benchmarks. No production usage figures are claimed.

How it runs

Proof first, then engineering.

The sequence is deliberate. Most failed AI projects were architected before anyone agreed what accurate meant.

  1. 01

    Define the failure that matters

    We start from the error the business cannot absorb, because that decides the architecture. Silent failure on a merged table cell is a different problem from a clumsy sentence.

  2. 02

    Build the evaluation first

    Ground truth is drawn from real material and validated by hand before tuning begins. Accuracy work then runs as measured campaigns rather than as ad hoc adjustment.

  3. 03

    Decide on evidence

    Where two approaches are plausible we run both behind a consensus check and retire the weaker one on the numbers, not on preference.

  4. 04

    Design the guardrails with the feature

    Residency, spend caps, rate limits, access control and the kill switch are part of the first release, not a hardening phase that gets deferred.

  5. 05

    Ship into a review surface

    A person stays in the loop where the record is consequential. The system earns autonomy over time, once the evaluation history supports it.

Deliverables

What you get,
and what you can defend.

An AI system creates obligations as well as capability. Both are part of the handover.

What goes live

A supported service, not a notebook.

  • Deployed infrastructure as code, with environments for UAT and production
  • An evaluation suite and its results, re-runnable on every change
  • Provenance and audit trails on every output that reaches a record
  • Cost ceilings enforced in policy, with a quantified worst case

What you can defend

The questions your risk and audit functions will ask, answered in writing.

  • Where data is stored, where inference runs and which region holds it
  • Which decisions a person confirms, and where the record of that sits
  • How accuracy was measured, on what sample, and what it is now
  • What happens when the model is wrong, and who is accountable for it
Technology

What we work in.

Model providers change. The engineering around them is where the durable work sits.

PythonTypeScriptLangGraphAWS BedrockRetrieval pipelinesDoclingNext.jsReactVector & hybrid searchSurrealDBPostgreSQLCloudFormationAWS EC2 GravitonCloudWatchSalesforce Connected AppsBitbucket PipelinesAgentforceModel evaluation harnesses

Bring us the bottleneck,
not the technology.

Describe the process that is absorbing people. We will tell you whether AI is the right answer, including when it is not.