Service detail

Enterprise AI Development

AI deployments that run on your own data, on your own server, with accuracy that is measured. Neither the prompt nor the answer crosses your boundary.

Done Dynamics builds enterprise AI deployments: on-premise large language models, internal assistants, document extraction pipelines and AI agents. We run the models on our own Mac Studio fleet and publish the builds we produce with their measured throughput and memory figures. The architecture we propose to clients is the one we use in our own operations every day.

What we deploy

Four kinds of deployment

All four share one trait: work a person already does today, repeatedly, in a way that can be measured.

Internal Assistant

An assistant that works on your own documents, records and procedures, cites the source of every answer, and stays inside the permissions of the person asking.

Document Extraction

Structured fields pulled from invoices, contracts, specifications and forms. A confidence value per field, consistency checks, and handover to a person where certainty is low.

Classification and Routing

Pipelines that route an inbound request to the right team, rank complaints by urgency and label free text — measurable accuracy, working alongside a rules engine.

AI Agents

Systems that act rather than only answer. Human approval on every irreversible operation, a deliberately narrow tool surface, and an audit record of who each action ran for.

2 minutes · 4 questions

Would an AI investment pay off in your operation?

Four questions. If your answers point to an off-the-shelf subscription, we say so.

Is there work your team repeats daily that runs on text or documents?

Question 1 / 4

Is there work your team repeats daily that runs on text or documents?

Why on premise

The answer should not live in the small print

Paste a contract into a hosted AI tool and that text leaves your building. An on-premise deployment removes the question entirely.

  • Data does not leave

    The model runs on your hardware or on our servers in Turkey. Neither prompt nor answer crosses that boundary, and we verify it together after setup.

  • Measured, not guessed

    We measured how fast each model runs and how much memory it wants on our own machines. The figures are public on the model cards we publish.

  • Start small, grow on numbers

    A four-week pilot aimed at one process. It ends in a measured accuracy figure rather than a demo, and the expansion decision follows that figure.

  • No provider lock-in

    We prefer openly licensed models. Deployment scripts, prompt templates and the evaluation set become yours by contract.

  • We run this ourselves

    A Mac Studio with 512 GB of unified memory and a second machine beside it. The architecture we propose is the one we use in our own operations daily.

  • GDPR and KVKK aligned

    The cross-border transfer question never opens. The process map, inventory entries and retention periods are prepared with the deployment.

Use cases

Where it pays for itself

From tender preparation to accounting documents, from institutional memory to field reports.

Tender and Specification Matching

Draft matching of specification clauses to your product catalogue — hours returned to the most experienced person in the room.

Accounting Document Pipeline

Field extraction from invoices and statements, automatic total and tax consistency checks, uncertain records routed to a person.

Institutional Memory

A retrieval layer over years of documents, email and meeting notes that answers with citations rather than assertions.

Customer Correspondence Drafts

Inbound messages classified and draft replies produced, with the decision to send always left to a person.

Quality and Non-conformance Analysis

Free-text field reports grouped into categories so that recurring root causes become visible.

Developer and IT Tooling

Log analysis, error classification and internal documentation assistants that sit on your own server.

Architectural decisions

Safety comes from architecture, not from the model

Most of the risk in an AI deployment is not what the model can do but what it is permitted to do. Those decisions are written down and visible.

  • Model choice follows the job; the reasoning and the measurements are shared in writing.
  • The quantisation decision follows the latency budget — four precisions of one model differ by up to threefold in speed and threefold in memory.
  • Retrieval runs under the asking identity; a document that person cannot see never enters the candidate pool.
  • Every answer cites its source, because an output that cannot be verified is a risk rather than information.
  • The confidence threshold is tunable: where the system is unsure it hands over instead of answering.
  • A held-out evaluation set measures accuracy regularly, and drift triggers intervention.
  • In agent deployments every irreversible action passes human approval and the tool surface stays narrow.
  • Content arriving from outside is marked as data — the defence against prompt injection comes from architecture, not from the model.
Measured data

We pick models on numbers, not impressions

We measured four precisions of the same model on the same machine. The gap is wide enough to decide whether a deployment is possible at all. Every build is public on our Hugging Face profile.

Build Generation Peak memory
4-bit37.9 tok/s15.5 GB
6-bit27.9 tok/s22.2 GB
8-bit22.2 tok/s28.9 GB
bf1612.7 tok/s54.1 GB

Mac Studio M3 Ultra, 512 GB unified memory. Single run, one prompt — an order-of-magnitude guide, not a benchmark.

Hosting

Where will the model live?

The first technical decision in an AI project is not which model but where it runs. There are three options, and the data itself decides which one is right.

On-Premise

The model files sit on a server in your own building. Neither prompt nor answer leaves the corporate network, and we verify together after setup that no outbound connection exists. The hardware investment is yours, and so is the audit trail.

Who it suits: Healthcare, legal, finance and defence — organisations where data leaving the building is ruled out from the start.

Private Hosting

A dedicated, unshared server on our hardware in Turkey. Resources are not split with anyone else, the network is separate, and access is limited to the people you name. The hardware cost spreads across a monthly service and maintenance stays with us.

Who it suits: Organisations that do not want the hardware burden of an on-premise deployment but cannot put their data on a public cloud.

Hybrid

Sensitive workloads on premise, the rest on private hosting. A different model can run on each side; the routing rule follows the data classification and produces an auditable record.

Who it suits: Where one organisation runs both processes handling personal data and processes producing public content.

We provide private hosting for workloads beyond AI too — application servers, databases and web hosting on the same infrastructure. Details on the server hosting and web hosting page.

Process

A measurable result in four weeks

Five stages, each with a visible output, ending in a number rather than a demo.

  1. 01

    Process Selection

    Which task, who does it, how many hours a day, what happens when it goes wrong. Without answers to those four we do not write a proposal.

  2. 02

    Samples and Baseline

    An evaluation set built from real documents, real questions and the answers people actually give today.

  3. 03

    Deployment

    Model and hardware are chosen, the pipeline stands up, and the first accuracy measurement is taken and reported.

  4. 04

    Interface and Permissions

    User interface, authentication and access control are added; source citation is wired into every answer.

  5. 05

    Pilot and Handover

    A small team uses it on real work and errors are logged. If the measured result holds, scope expands; handover training is delivered.

Commercial model

Pilot first, contract second

Proposals that start broad tend to still be demos six months later. So we start with a fixed-scope pilot: one process, four weeks, a measured accuracy figure at the end. If the result holds, the annual maintenance and expansion contract is built on that number — scope and price written up front and unchanged through the year. If it does not, where it breaks is visible, and the decision to stop is yours.

Ready for your next software project?

Book a free 30-minute discovery call with our team.

Certifications

Our network and cyber security work is carried out by a team holding internationally recognised Cisco certification.

Cisco CyberOps Associate badge

Cisco CyberOps Associate

Issued by Cisco · Holder: Devrim Tunçer

A certification covering security operations centre (SOC) competency: security monitoring, incident response and analysis of network attacks. It is the foundation we rely on for intrusion detection, log correlation and post-incident response work.

Cisco CCNA Training

expired

Cisco training certificate · completed January 2023

Covers networking fundamentals: routing, switching, IP addressing and network security. The knowledge base we draw on for enterprise network setup and segmentation.

Enterprise AI development: where does your data sit?

The question that decides the architecture

The moment you paste a contract into a hosted AI tool, that text leaves your building. Where it goes, how long it is kept and whether it is used for training are answers buried in the small print. Enterprise AI development exists to remove that question entirely: the model runs on your own server or on our hardware in Türkiye, and neither the prompt nor the answer crosses that boundary.

We describe this as a deployment we actually run, not a slide. A Mac Studio with 512 GB of unified memory and a smaller machine beside it, the models running on them, and the measured throughput of those models. Because we can stand the same setup up on the client side, keeping data inside is a starting point here rather than a constraint.

Which work belongs to AI, and which does not

Any task whose rule can be written belongs to ordinary software. Summing an invoice, raising an alert when stock drops below a threshold, counting days between two dates — none of these need a model, and using one makes them slower and less reliable. AI pays where the rule cannot be written in advance: understanding what a text says, matching the same fact written two different ways, extracting structured fields from free text, summarising a long document.

The three families we see most often are document extraction, internal search across years of accumulated files, and classification and routing of inbound requests. What they share is that a person is already doing the work today — which means the gain is measurable rather than hypothetical: how many people, how many hours a day. As a corporate software company, that is the number we ask for in the first call.

Model size, quantisation and the latency budget

Model selection usually starts with which one is best, when the useful question is what the smallest good-enough model for this job is. We measured four precisions of the same model on the same machine: the most compressed build produced roughly 38 tokens per second while the uncompressed one managed 13. On memory the gap is sharper still — 15.5 GB against 54 GB. Same model, same hardware, three times the speed and a third of the memory.

The practical rule that follows: write down your latency budget first. When a user waits at the screen, the difference between two seconds and ten decides whether the product gets used. In an overnight batch job nobody cares, and the choice there goes to quality. Running two different models for two different jobs inside one organisation is ordinary design, not an exception.

Connecting to internal knowledge

A model does not need retraining to know your company. Instead your documents are made searchable, the relevant passages are retrieved when a question arrives, and the model is given those passages to answer from. The quality of that setup is decided on the retrieval side far more than the model side — in most failed deployments we have seen, the model was fine and the search was not.

Two things are non-negotiable. Every answer must cite the document and section it came from, because an assistant that cannot be checked stops being used after a few wrong answers. And retrieval must run under the identity of the person asking, so that an employee without access to the finance folder cannot reach its contents by asking a question. Both are built in on day one, never bolted on later.

A pilot that produces a number

The most common failure is starting broad. An assistant covering every department is still a demo six months later, while a pilot aimed at one process is in front of real users in week four. Week one we watch the process and collect real examples; week two the deployment stands up and accuracy is measured; week three the interface and authorisation are added; week four a small team uses it on real work and errors are logged.

What you hold at the end is not a demo but a number: in how many of a hundred cases did the system get it right, and how often did a human have to correct it. If that number is good enough, expanding is an easy decision. If it is not, where it breaks is visible and cheap to fix.

Frequently asked questions

Does our data really stay inside?
With an on-premise deployment the model files sit on your hardware and neither prompt nor answer leaves that machine. We verify together after setup that no outbound connection exists. Where we host it instead, the hardware is in Türkiye and the data processing agreement states so in writing.
Which model do you use?
The choice follows the job and the reasoning is shared in writing. We prefer openly licensed models because they can run on premise and can be handed over. The measured speed and memory figures for the builds we publish are public.
Can it connect to our ERP or document system?
Yes, to the extent the system allows. If it exposes an interface we use it; otherwise we work at the database level or through export files. We check this in the first call, because it usually decides the timeline.
How is the risk of a wrong answer managed?
Three layers: every answer cites its source, a confidence threshold hands uncertain cases to a person, and a held-out evaluation set is re-run regularly to catch drift. We do not promise zero errors — we make errors visible and measurable.
How soon do we see something usable?
Four weeks for a pilot aimed at one process. At the end of it you have a deployment real users have tried on real work, and a measured accuracy figure.
Can our own team take it over?
Yes, that is the goal. Deployment scripts, prompt templates and the evaluation set are handed over, with a handover session for your team. Us continuing the maintenance is an option, not a requirement.