Production is the hard part.

We build AI systems, workflow automation and operational software that work beyond the demo. Integrations, permissions, evaluation, failure handling, observability and handover are designed in rather than retrofitted.

What we build

Five practices, one way of working. Whatever the work is, it is built to survive contact with production, and measured before you depend on it.

  1. Web and commerce platformsMarketing sites, customer portals and headless commerce that stay fast on a phone connection.What is included

    At handover: a deployed site or storefront, its content model, and the pipeline that ships changes to it

    Try a working demo
  2. Mobile and connected productsiOS and Android from one codebase, and the hardware that reports back to them.What is included

    At handover: signed iOS and Android builds, both store listings, and an automated release pipeline

    Try a working demo
  3. Enterprise operationsThe systems operations actually run on: stock, orders, logistics, and the internal tools around them.What is included

    At handover: connected workflows, role-based access, audit trails, and the dashboards operations run on

    Try a working demo
  4. Data and cloud deliveryPipelines, warehousing, access control, and the release machinery that puts a change live safely.What is included

    At handover: data pipelines, a warehouse, access controls, and deployments with monitoring attached

    Try a working demo
  5. AI and automationAssistants that cite their source, agents with a job description, document intelligence.What is included

    At handover: an evaluation set, the accuracy scores it produces, and the service deployed behind an interface

    Try a working demo

All 15 services in detail What each one includes, and what it does not.

Systems we have built

Named by system rather than by client, because most of it is covered by an agreement we intend to keep. No percentages: what each one carries is the workflow, the engineering decisions, and the operating change that followed.

  1. Logistics · Workflow automation · Systems integration

    Enterprise operations

    Freight Operations Platform

    The shipment lifecycle ran across disconnected channels, so information arrived inconsistently and nobody owned the current status.

    A web operational system structured around the shipment lifecycle, from quotation intake through to completion, designed to talk to carrier systems over an API.

    What changed

    Quotation, booking, tracking and completion became one traceable flow rather than a sequence of unrelated administrative tasks.

  2. Architecture · 3D visualization · WebGL · Gesture interaction

    Mobile and connected products

    Spatial Architecture Visualization System

    Architects spend hours producing renders, and clients still cannot read a plan as a sense of space.

    A system design taking an uploaded 2D floor plan through to an interactive 3D scene, shown on a reflective display several people can gather round without a headset.

    What changed

    The design closes the interpretation gap without a headset, which was the constraint the whole architecture was built around.

  3. AI automation · Workflow orchestration · LynchPine product

    AI and automation

    WorkShip

    Every automation engagement rebuilt the same mechanics: state, retries, tool execution, decision points, model calls, observability and cost.

    A workflow-driven automation harness where the environment used to develop and test a workflow is the environment that runs it in production.

    What changed

    Orchestration stops being rebuilt per engagement, so implementation can start on the customer's workflow instead of the plumbing under it.

All the work

One yes, ten jobs

Deciding to build something is the small decision. What follows is the part that fills a year.

Say yes to a customer portal, an internal tool or an assistant that answers questions about your own documents, and you have taken on the work behind it. Somebody has to choose the stack and defend the choice in two years. Somebody has to decide what the system is allowed to touch, and prove it to a client's security questionnaire. Somebody has to keep it up on the Monday the warehouse is behind, watch the bill, answer to a model provider that just retired the version you depend on, and make sure people outside the company can find the thing at all.

Few companies are short of ideas. They are short of the layer underneath: the engineers, the tooling, and the habit of running software that other people's work depends on.

Hiring for skills that keep moving
The engineer who can build a retrieval system, run a release pipeline and read an evaluation set is the engineer every company is trying to hire this year. Job ads for one person who does all three are how projects get stuck before they start.
Security that has to hold up to somebody else's audit
Access control, secrets, dependencies, audit trails. It is quiet work until a customer sends a security questionnaire, and then it is the only work.
Being found, by people and by assistants
Search still sends the visitors. Increasingly an assistant reads a page and answers on your behalf, and a site it cannot parse is a site it does not quote.
Keeping it alive after launch
Someone watches the logs, patches what needs patching, and renews the certificate before it expires. This is the part that gets dropped first and noticed last.
Knowing when the answer is no
The most expensive systems are the ones nobody could talk the business out of. A team paid to build rarely argues for a form and a query instead.

Start from the problem instead Seven weeks that are not working, and what we do about each.

Projects fail at the joins

A demo needs one good answer. A production system needs the ten thousandth, at the end of a bad week, from someone who filled the form in wrong. The engineering between those two points is where projects stall: where the data comes from, what the system is allowed to touch, what happens when a run fails, and how you know six months later that it still works.

Those are the joins. They rarely show up in a demo, and they are most of the work.

agent playground

Pick a request below. It runs here.

A support system, running. Each request wakes the agents it needs, touches only what it is allowed to touch, and writes down what it did. The third one is refused.

Task: What are our termination rights on the supplier agreement?

run log

  1. Pick a request above. The run is drawn on the left and written here as it happens.

Take the queue yourself Ninety seconds, four agents, one budget.

What a system we build looks like

Five stages, and a loop. The last one feeds the second, which is the difference between a system that degrades quietly and one that gets better on purpose.

  1. Your data, not the internet's

    Contracts, tickets, SOPs, the spreadsheet one person maintains. We map it before deciding what a model may see.

  2. Retrieval that cites its source

    Hybrid keyword and vector search with reranking, tuned on your questions instead of a public benchmark.

  3. The part that holds

    Routing, tool access, shared state, handoffs. Where most AI projects quietly come apart.

  4. Agents with a job description

    Scoped and permissioned, so approving one is a decision someone can actually make.

  5. Proof it still works

    A test set from real queries, run on every change, then fed back so the system improves on purpose.

sources seen from above

How it finds the right paragraph

Ask a question of a thousand documents and this is what happens. Your material is placed by meaning rather than by folder; the question lands in that space, and the passages nearest it come back, each one still carrying the document it came from. That last part is why an answer can be checked.

retrieval

Keep scrolling to run the search.

  • a chunk of your material
  • how far the search reached
  • what it returned

You be the retrieval step

Here is a question and six passages from a fictional company's documents. Choose up to three for the model to read, then see what it answers. The model is the same every time. Only what you hand it changes.

The questionWhat is the refund window for enterprise customers?

Passages to hand the model0 of 3 chosen

Where we differ

  1. Most of what gets called AI should be an if-statement

    A model is a poor substitute for deterministic logic and roughly a thousand times the price. We put one in the path only where the work genuinely calls for judgement. If your problem is really a rules engine, we will say so, and the bill will be smaller for it.

    What a decision costs

    Per thousand decisions, at the point in the path where the choice is actually made.

    USD per 1,000 decisions

    1. A branch in code

      0.002

      Deterministic, testable, and the same answer every time.

    2. A small model

      0.15

      Worth it where the input is genuinely unstructured.

    3. A frontier model

      2.00

      Worth it where the work genuinely calls for judgement.

    IllustrativeRounded hard on purpose: the argument is the order of magnitude, not the third significant figure. Linear scale, so the first bar is a hairline, which is the point of it. Published rates move every few months; the ratio between a branch and a model call has not.

  2. A demo is not evidence

    Anything can be made to work once, on a question the vendor picked. We would rather build an evaluation set from your own queries first and show you how the system scores against it. That number is what you can take to a board.

    The demo always wins

    Two lines, eight builds. One is the query chosen for the demo. The other is a held-out set of real questions. Only one of them is evidence.

    • The query picked for the demo
    • A held-out set of your questions
    Score against each source, by build. Illustrative figures.
    BuildThe query picked for the demoA held-out set of your questions
    0196%47%
    0297%55%
    0396%58%
    0498%69%
    0597%74%
    0698%79%
    0798%84%
    0899%86%

    IllustrativeThe shape is what matters: a cherry-picked query is at the top from the first build and stays there, so it can tell you nothing about whether the system improved. The line that can is the one that does not reach it.

  3. Nobody should marry a model provider

    The frontier moves every few months. We put the model behind an interface, so switching one out is a config change and a couple of eval runs. When a provider raises prices or retires the version you depend on, you get to decide what happens next, on your own timetable.

    Swapping the model

    The system talks to an interface, never to a provider. Changing the model behind it is a config change, and the evaluation set runs before the new one takes any traffic.

    A schematic of the arrangement, not a comparison of providers.

Where the work goes

Based in Colombo, working with clients in Australia, Canada, India, Norway, Sri Lanka and the United Kingdom. Most of it is remote, and the timezone question is usually the real one, so the offsets are on the map.

Where the working days meet

  1. AustraliaUTC +10:00

    Working day runs 04:30 to 13:30 Colombo time.

    4.5 h shared

  2. IndiaUTC +05:30

    Working day runs 09:00 to 18:00 Colombo time.

    9 h shared

  3. Sri Lankathe studio

    Working day runs 09:00 to 18:00 Colombo time.

    9 h shared

  4. NorwayUTC +01:00

    Working day runs 13:30 to 22:30 Colombo time.

    4.5 h shared

  5. United KingdomUTC +00:00

    Working day runs 14:30 to 23:30 Colombo time.

    3.5 h shared

  6. CanadaUTC -05:00

    Working day runs 19:30 to 04:30 Colombo time.

    0 h shared

Colombo working day. Each country's 09:00 to 18:00 working day, placed on Colombo's clock. The offsets are standard time, so daylight saving moves some rows by an hour for part of the year.

What happens after you email us

Four phases, whatever the work is. The point of running it this way is to force the decisions into the open early, because the ones nobody makes are the ones an AI tool will quietly make for you, in code, where they are expensive to find.

  1. ClarifyWeek 1

    We find the decisions nobody has made yet, the ones an AI tool would otherwise make for you, silently, in code. You get a written scope and a fixed price before anything is built.

  2. PlanWeeks 2-3

    Architecture, data flow, agent boundaries, and what counts as working. All of it agreed while it is still cheap to argue about, and written down so nobody relitigates it in week six.

  3. Build and verifyWeeks 3-7

    Small changes, each one run against the evaluation set. You get a live environment in the first week and watch it fill in, so nothing arrives as a surprise at handover.

  4. Learn and adjustWeek 8 onward

    Real use produces cases the plan missed. Those become new tests, the system improves against them, and the loop keeps running after handover, with your team or with ours.

Payments

  1. to commence
  2. on design approval
  3. on completion
A typical six to eight week build. The dates and the price for yours are in the proposal before anything starts.

The record

20+ projects delivered across 5+ industries, and a typical build runs 6-8 weeks.

Built by the team that runs it

A studio hands over a repository at the end. A managed provider runs something it did not build. We do both halves, and you pick where the line sits.

The same engineers who design the system write the pipeline that ships it, the tests that gate it and the runbook the next team reads. When something breaks at nine on a Tuesday, nobody is reading the code for the first time.

You decide what happens at handover. Take all of it, keys and accounts included, and run it yourself. Leave it with us on a monthly arrangement. Or split it, with your team owning the product and ours keeping the delivery and evaluation machinery running.

  • We build it, you run it

    The default. Code, prompts, evaluation sets, infrastructure and accounts transfer to you on final payment, with documentation and a training session so your team can take it from there.

  • We build it and stay

    A monthly arrangement for systems in production: evaluation runs on a schedule, prompts and configs move as models move, and a named response time when something is wrong.

  • You have a team, we fill the gap

    Your engineers own the product and we cover the layer they do not have yet, whether that is retrieval, the release pipeline or the evaluation harness. We write it so your team can read it.

Why a studio rather than a hire The long version, with what you can check.

What you can hold us to

Evaluation before opinion
Every AI system we ship has a test set drawn from your real queries and a score you can track over time. When we claim something got better, the number comes with it.
You own all of it
Source code, prompts, evaluation sets, infrastructure and accounts transfer to you on final payment. Nothing stays licensed from us. If you would rather we kept running it afterwards, that is a separate arrangement and it does not change who owns any of it.
No model lock-in
The model sits behind an interface. Swapping it is a config change and an eval run, so you are never stuck with a provider because leaving would cost too much.
We will talk you out of it
If your problem is better solved by a form, a query or a rule, we will say so before you have paid for a model to do it worse.
Fixed scope, fixed price
You approve a written scope and a number before work starts. Changes get quoted, not absorbed silently and billed later.
Built to be handed over
Every project ships with documentation and a training session. If you replace us next year, the next team can read what we left.

Bring the problem, not a specification

Send a paragraph about the problem. We reply within two working days with questions, a rough shape, and an honest answer on whether we are the right people for it.