Production is the hard part.
We build AI systems, workflow automation and operational software that work beyond the demo. Integrations, permissions, evaluation, failure handling, observability and handover are designed in rather than retrofitted.
What we build
Five practices, one way of working. Whatever the work is, it is built to survive contact with production, and measured before you depend on it.
- Web and commerce platformsMarketing sites, customer portals and headless commerce that stay fast on a phone connection.What is included
At handover: a deployed site or storefront, its content model, and the pipeline that ships changes to it
Try a working demo - Mobile and connected productsiOS and Android from one codebase, and the hardware that reports back to them.What is included
At handover: signed iOS and Android builds, both store listings, and an automated release pipeline
Try a working demo - Enterprise operationsThe systems operations actually run on: stock, orders, logistics, and the internal tools around them.What is included
At handover: connected workflows, role-based access, audit trails, and the dashboards operations run on
Try a working demo - Data and cloud deliveryPipelines, warehousing, access control, and the release machinery that puts a change live safely.What is included
At handover: data pipelines, a warehouse, access controls, and deployments with monitoring attached
Try a working demo - AI and automationAssistants that cite their source, agents with a job description, document intelligence.What is included
At handover: an evaluation set, the accuracy scores it produces, and the service deployed behind an interface
Try a working demo
All 15 services in detail What each one includes, and what it does not.
Systems we have built
Named by system rather than by client, because most of it is covered by an agreement we intend to keep. No percentages: what each one carries is the workflow, the engineering decisions, and the operating change that followed.
Freight Operations Platform
The shipment lifecycle ran across disconnected channels, so information arrived inconsistently and nobody owned the current status.
A web operational system structured around the shipment lifecycle, from quotation intake through to completion, designed to talk to carrier systems over an API.
What changed
Quotation, booking, tracking and completion became one traceable flow rather than a sequence of unrelated administrative tasks.
Spatial Architecture Visualization System
Architects spend hours producing renders, and clients still cannot read a plan as a sense of space.
A system design taking an uploaded 2D floor plan through to an interactive 3D scene, shown on a reflective display several people can gather round without a headset.
What changed
The design closes the interpretation gap without a headset, which was the constraint the whole architecture was built around.
WorkShip
Every automation engagement rebuilt the same mechanics: state, retries, tool execution, decision points, model calls, observability and cost.
A workflow-driven automation harness where the environment used to develop and test a workflow is the environment that runs it in production.
What changed
Orchestration stops being rebuilt per engagement, so implementation can start on the customer's workflow instead of the plumbing under it.
One yes, ten jobs
Deciding to build something is the small decision. What follows is the part that fills a year.
Say yes to a customer portal, an internal tool or an assistant that answers questions about your own documents, and you have taken on the work behind it. Somebody has to choose the stack and defend the choice in two years. Somebody has to decide what the system is allowed to touch, and prove it to a client's security questionnaire. Somebody has to keep it up on the Monday the warehouse is behind, watch the bill, answer to a model provider that just retired the version you depend on, and make sure people outside the company can find the thing at all.
Few companies are short of ideas. They are short of the layer underneath: the engineers, the tooling, and the habit of running software that other people's work depends on.
- Hiring for skills that keep moving
- The engineer who can build a retrieval system, run a release pipeline and read an evaluation set is the engineer every company is trying to hire this year. Job ads for one person who does all three are how projects get stuck before they start.
- Security that has to hold up to somebody else's audit
- Access control, secrets, dependencies, audit trails. It is quiet work until a customer sends a security questionnaire, and then it is the only work.
- Being found, by people and by assistants
- Search still sends the visitors. Increasingly an assistant reads a page and answers on your behalf, and a site it cannot parse is a site it does not quote.
- Keeping it alive after launch
- Someone watches the logs, patches what needs patching, and renews the certificate before it expires. This is the part that gets dropped first and noticed last.
- Knowing when the answer is no
- The most expensive systems are the ones nobody could talk the business out of. A team paid to build rarely argues for a form and a query instead.
Start from the problem instead Seven weeks that are not working, and what we do about each.
Projects fail at the joins
A demo needs one good answer. A production system needs the ten thousandth, at the end of a bad week, from someone who filled the form in wrong. The engineering between those two points is where projects stall: where the data comes from, what the system is allowed to touch, what happens when a run fails, and how you know six months later that it still works.
Those are the joins. They rarely show up in a demo, and they are most of the work.
agent playground
Pick a request below. It runs here.
A support system, running. Each request wakes the agents it needs, touches only what it is allowed to touch, and writes down what it did. The third one is refused.
Task: What are our termination rights on the supplier agreement?
run log
- Pick a request above. The run is drawn on the left and written here as it happens.
Take the queue yourself Ninety seconds, four agents, one budget.
What a system we build looks like
Five stages, and a loop. The last one feeds the second, which is the difference between a system that degrades quietly and one that gets better on purpose.
Your data, not the internet's
Contracts, tickets, SOPs, the spreadsheet one person maintains. We map it before deciding what a model may see.
Retrieval that cites its source
Hybrid keyword and vector search with reranking, tuned on your questions instead of a public benchmark.
The part that holds
Routing, tool access, shared state, handoffs. Where most AI projects quietly come apart.
Agents with a job description
Scoped and permissioned, so approving one is a decision someone can actually make.
Proof it still works
A test set from real queries, run on every change, then fed back so the system improves on purpose.
How it finds the right paragraph
Ask a question of a thousand documents and this is what happens. Your material is placed by meaning rather than by folder; the question lands in that space, and the passages nearest it come back, each one still carrying the document it came from. That last part is why an answer can be checked.
retrieval
Keep scrolling to run the search.
- a chunk of your material
- how far the search reached
- what it returned
You be the retrieval step
Here is a question and six passages from a fictional company's documents. Choose up to three for the model to read, then see what it answers. The model is the same every time. Only what you hand it changes.
The questionWhat is the refund window for enterprise customers?
Where we differ
Most of what gets called AI should be an if-statement
A model is a poor substitute for deterministic logic and roughly a thousand times the price. We put one in the path only where the work genuinely calls for judgement. If your problem is really a rules engine, we will say so, and the bill will be smaller for it.
What a decision costs
Per thousand decisions, at the point in the path where the choice is actually made.
USD per 1,000 decisions
IllustrativeRounded hard on purpose: the argument is the order of magnitude, not the third significant figure. Linear scale, so the first bar is a hairline, which is the point of it. Published rates move every few months; the ratio between a branch and a model call has not.
A demo is not evidence
Anything can be made to work once, on a question the vendor picked. We would rather build an evaluation set from your own queries first and show you how the system scores against it. That number is what you can take to a board.
The demo always wins
Two lines, eight builds. One is the query chosen for the demo. The other is a held-out set of real questions. Only one of them is evidence.
- The query picked for the demo
- A held-out set of your questions
Score against each source, by build. Illustrative figures. Build The query picked for the demo A held-out set of your questions 01 96% 47% 02 97% 55% 03 96% 58% 04 98% 69% 05 97% 74% 06 98% 79% 07 98% 84% 08 99% 86% IllustrativeThe shape is what matters: a cherry-picked query is at the top from the first build and stays there, so it can tell you nothing about whether the system improved. The line that can is the one that does not reach it.
Nobody should marry a model provider
The frontier moves every few months. We put the model behind an interface, so switching one out is a config change and a couple of eval runs. When a provider raises prices or retires the version you depend on, you get to decide what happens next, on your own timetable.
Swapping the model
The system talks to an interface, never to a provider. Changing the model behind it is a config change, and the evaluation set runs before the new one takes any traffic.
A schematic of the arrangement, not a comparison of providers.
Where the work goes
Based in Colombo, working with clients in Australia, Canada, India, Norway, Sri Lanka and the United Kingdom. Most of it is remote, and the timezone question is usually the real one, so the offsets are on the map.
Where the working days meet
AustraliaUTC +10:00
Working day runs 04:30 to 13:30 Colombo time.4.5 h shared
IndiaUTC +05:30
Working day runs 09:00 to 18:00 Colombo time.9 h shared
Sri Lankathe studio
Working day runs 09:00 to 18:00 Colombo time.9 h shared
NorwayUTC +01:00
Working day runs 13:30 to 22:30 Colombo time.4.5 h shared
United KingdomUTC +00:00
Working day runs 14:30 to 23:30 Colombo time.3.5 h shared
CanadaUTC -05:00
Working day runs 19:30 to 04:30 Colombo time.0 h shared
What happens after you email us
Four phases, whatever the work is. The point of running it this way is to force the decisions into the open early, because the ones nobody makes are the ones an AI tool will quietly make for you, in code, where they are expensive to find.
- ClarifyWeek 1
We find the decisions nobody has made yet, the ones an AI tool would otherwise make for you, silently, in code. You get a written scope and a fixed price before anything is built.
- PlanWeeks 2-3
Architecture, data flow, agent boundaries, and what counts as working. All of it agreed while it is still cheap to argue about, and written down so nobody relitigates it in week six.
- Build and verifyWeeks 3-7
Small changes, each one run against the evaluation set. You get a live environment in the first week and watch it fill in, so nothing arrives as a surprise at handover.
- Learn and adjustWeek 8 onward
Real use produces cases the plan missed. Those become new tests, the system improves against them, and the loop keeps running after handover, with your team or with ours.
Payments
- 40%to commence
- 30%on design approval
- 30%on completion
The record
20+ projects delivered across 5+ industries, and a typical build runs 6-8 weeks.
Built by the team that runs it
A studio hands over a repository at the end. A managed provider runs something it did not build. We do both halves, and you pick where the line sits.
The same engineers who design the system write the pipeline that ships it, the tests that gate it and the runbook the next team reads. When something breaks at nine on a Tuesday, nobody is reading the code for the first time.
You decide what happens at handover. Take all of it, keys and accounts included, and run it yourself. Leave it with us on a monthly arrangement. Or split it, with your team owning the product and ours keeping the delivery and evaluation machinery running.
We build it, you run it
The default. Code, prompts, evaluation sets, infrastructure and accounts transfer to you on final payment, with documentation and a training session so your team can take it from there.
We build it and stay
A monthly arrangement for systems in production: evaluation runs on a schedule, prompts and configs move as models move, and a named response time when something is wrong.
You have a team, we fill the gap
Your engineers own the product and we cover the layer they do not have yet, whether that is retrieval, the release pipeline or the evaluation harness. We write it so your team can read it.
Why a studio rather than a hire The long version, with what you can check.
What you can hold us to
- Evaluation before opinion
- Every AI system we ship has a test set drawn from your real queries and a score you can track over time. When we claim something got better, the number comes with it.
- You own all of it
- Source code, prompts, evaluation sets, infrastructure and accounts transfer to you on final payment. Nothing stays licensed from us. If you would rather we kept running it afterwards, that is a separate arrangement and it does not change who owns any of it.
- No model lock-in
- The model sits behind an interface. Swapping it is a config change and an eval run, so you are never stuck with a provider because leaving would cost too much.
- We will talk you out of it
- If your problem is better solved by a form, a query or a rule, we will say so before you have paid for a model to do it worse.
- Fixed scope, fixed price
- You approve a written scope and a number before work starts. Changes get quoted, not absorbed silently and billed later.
- Built to be handed over
- Every project ships with documentation and a training session. If you replace us next year, the next team can read what we left.
Bring the problem, not a specification
Send a paragraph about the problem. We reply within two working days with questions, a rough shape, and an honest answer on whether we are the right people for it.