Depth 10 · Service line

AI engineering and automation

AI engineering and automation is the work of putting language models and automated workflows into a company's daily operations as maintained software, with tests, access controls and an owner. It covers assistants that answer from your own documents, automation of repetitive back office steps, the checks that catch a model getting worse, and preparation for the rules that govern AI in the European market.

Inside this line

Four sub-services: assistants, process automation, model operations and governance readiness under the EU AI Act.

Input
A named business process or document set, the people who do that work today, and access to the systems the work touches.
Method
We describe the current process step by step before choosing any model. Then we build the smallest version that does one job, test it against real examples your team has checked, and widen it only once the tests hold.
Deliverable
Working software in your environment, an evaluation set your team can rerun, and written documentation of what the system does, what it cannot do and who is responsible for it.
Measured by
Accuracy against a reviewed test set, the share of outputs a person had to correct, and time spent on the process before and after, measured by your team on your own data.

Level 3 · 01

Assistants and copilots

An AI assistant, sometimes called a copilot, is a tool that answers questions or drafts text by combining a language model with a company's own information. A good one quotes the document it relied on, so the person reading can check it. It does not know anything outside the material it has been given, and it can still misread that material.

Inside it

  • Document ingestion that keeps access permissions
  • Retrieval with cited sources
  • Refusal when no source is found
  • Feedback capture from users
  • A written list of known limits

Measured by The share of test questions answered correctly with the right source cited, and how often the assistant correctly says it does not know.

Deliverable · Level 4→ Retrieval assistant over internal docsRetrieval assistant over internal docs: a working assistant, the test question set and a page of known limits written for its users.

Level 3 · 02

Process automation

Process automation means handing the repeatable steps of a business process to software, so that people handle the exceptions instead of every case. Some steps use a language model, for example reading an invoice or sorting an email. Most steps are ordinary code with rules. Automating a broken process does not repair it; it just makes it fail faster.

Inside it

  • Process mapping with the people doing the work
  • Removing steps before automating them
  • Rule based steps and model based steps kept separate
  • Human review queue for uncertain cases
  • Exception runbook

Measured by Cases handled without manual correction, cases sent to a person and why, and handling time, compared with the baseline your team records before the change.

Deliverable · Level 4→ Back office workflow rebuildBack office workflow rebuild: the redesigned process, the automation running in your systems, and a runbook for the exceptions.

Level 3 · 03

Model operations

Model operations is the routine that keeps an AI system working after launch. Model providers update their models, documents change and the questions people ask drift, and any of these can quietly lower quality. Model operations catches that with repeatable tests, logs and a clear process for changing prompts or models.

Inside it

  • Reviewed evaluation set from real cases
  • Automated scoring with human spot checks
  • Regression run on every prompt or model change
  • Production logging and failure review
  • Version history of prompts and models

Measured by Scores on the regression set over time, and the number of changes that were blocked before release because a score dropped.

Deliverable · Level 4→ Evaluation harness and regression setEvaluation harness and regression set: the test cases, the scoring rules and a script that runs them before any change ships.

Level 3 · 04

AI governance readiness

AI governance readiness is the preparation a company needs to meet the EU Artificial Intelligence Act, Regulation (EU) 2024/1689. The Act applies to companies that build AI systems or use them in the European market, including companies based outside the EU. It sorts AI systems by risk: some uses are banned, high risk uses carry duties such as risk management, documentation, human oversight and logging, and some systems must tell people they are dealing with AI.

Inside it

  • AI system inventory
  • Provider or deployer role per system
  • Risk classification against Regulation (EU) 2024/1689
  • Transparency and human oversight checks
  • Gap list with owners

Measured by Each system in the inventory has a recorded category, role and owner, and each gap is closed or carries a dated plan.

Deliverable · Level 4→ EU AI Act gap reviewEU AI Act gap review: the system inventory, a proposed risk category for each system with the reasoning, and a list of gaps, each with an owner.

Where this applies

Who this is for: Product and operations leads with an AI feature or a manual process they can describe step by step, and documents or data they own.

Sectors this line is set up for, not a list of clients.

Professional servicesIND 03
Assistants that answer from approved precedents and policies and show the passage they used.
LogisticsIND 04
Document intake and exception handling moved from inboxes into tracked workflows.
B2B softwareIND 02
Support copilots grounded in your own documentation, with an evaluation set run before each release.
Companies facing European regulation
A system inventory and risk classification as preparation for the EU AI Act.

Shape of the work

01Question02Retrieve passages03Draft with citations04Evaluate
Retrieval loop: a question is embedded, matching passages are retrieved from your documents, the model drafts an answer with citations, an evaluation step checks it, and failures feed back into the test set.

What you get

  • 10.1 · Assistants and copilotsRetrieval assistant over internal docs: a working assistant, the test question set and a page of known limits written for its users.
  • 10.2 · Process automationBack office workflow rebuild: the redesigned process, the automation running in your systems, and a runbook for the exceptions.
  • 10.3 · Model operationsEvaluation harness and regression set: the test cases, the scoring rules and a script that runs them before any change ships.
  • 10.4 · AI governance readinessEU AI Act gap review: the system inventory, a proposed risk category for each system with the reasoning, and a list of gaps, each with an owner.

What we work with

Languages and runtimes
TypeScript · JavaScript · Node · Python · Go · SQL · Bash
Back end and APIs
REST · GraphQL · webhooks · background workers · queues
Data
PostgreSQL · pgvector · Supabase · Redis · dbt · BigQuery
AI
Anthropic API · OpenAI API · open weight models on hosted inference · embeddings and vector search · retrieval over your own content · evaluation harnesses · prompt and model version control
CRM and messaging
HubSpot · Salesforce · transactional mail providers
Automation
n8n · scheduled jobs · event driven triggers

These are the tools we work with, not partnerships, certifications or resale agreements.

How this fits the other lines

When this is not the right line

If the process itself is unclear or disputed, automating it makes the disagreement faster, so fix the process first. If you need an answer that is always correct, a language model is the wrong tool: it can cite sources but it cannot guarantee truth.

Questions about this line

Which models do you use?
The model is chosen per task against an evaluation set built from your own examples. It can be an Anthropic or OpenAI model through the API, or an open weight model on hosted inference. The choice and its test results are written down so it can be changed later.
Will an AI assistant give wrong answers?
Sometimes, yes. Language models can produce fluent text that is false, so we design for checking: answers show the source passage they came from, the assistant says when it found nothing, and a regression test set flags when quality drops after a change.
Does our data leave our environment or get used to train someone else's model?
That depends on the model provider and the contract you hold with them, and we settle it with you before any data moves. Where it matters, we use providers whose business terms exclude training on your inputs, or models that run inside your own cloud account.
Which processes are worth automating first?
Ones that are already stable, happen often, follow written rules and have an output a person can check quickly. If a process is unclear or keeps changing, we fix the process first, because automating it only makes the same failures happen faster.

Name the process or the AI feature

Describe one process or one AI feature and what a correct result looks like. We reply with how we would test it before building anything.

Describe the process

An email exchange first, no call needed.