AI & Marketing · Consulting
Private & On-Premise AI: Models That Run on Your Hardware and Nowhere Else
If the data cannot leave the building, the model has to come to it.
The problem
What this usually looks like
Some businesses cannot send their work to a hosted API and already know it. Customer lists, pricing, signed contracts, medical and legal files, drawings under NDA, none of that belongs with a vendor whose terms change annually and whose retention policy you do not control. So one of two things happens: staff quietly paste it into a consumer chatbot anyway and the exposure occurs without anyone logging it, or AI gets banned outright and the productivity goes to a competitor who found a way. The third option is running the model on hardware you own, in a room you control.
How I approach it
I size the deployment to the actual workload rather than to a leaderboard. That means choosing an open-weight model that is good enough for the specific task, quantizing it to fit hardware you can justify buying, and measuring real throughput, tokens per second, concurrent users, time to first token. Before anything is promised to the business. llama.cpp and comparable runtimes make this practical on a single workstation-class GPU for most business workloads, and exposing an OpenAI-compatible endpoint means existing tools connect without being rewritten. The build finishes with documentation and handover so your own IT staff can run it, restart it and swap models without calling me.
Scope
What I actually do
Assess which workloads genuinely require on-premise inference and which are perfectly safe on a hosted API.
Select open-weight models sized to the task rather than to a benchmark score.
Quantize and test candidate models so quality loss is measured against your own examples, not assumed.
Size hardware honestly: VRAM, system memory, CPU offload, and say plainly what a given budget will and will not run.
Deploy with llama.cpp and similar runtimes behind an OpenAI-compatible endpoint, then benchmark throughput under realistic concurrency.
Build the cost comparison against per-token API billing at your projected volume, including hardware amortisation and power.
Set up retrieval over internal documents so the private model can answer from company files without anything leaving the network.
Document the stack, monitoring and restart procedure and hand it over to internal IT.
Deliverables
What you get
- Workload and data-sensitivity assessment separating what must stay in-house from what does not
- Model selection and quantization report with measured quality and speed on your own examples
- Hardware specification and sizing recommendation tied to a stated budget
- Deployed inference server with an OpenAI-compatible API endpoint
- Throughput benchmarks: tokens per second, concurrency limits, time to first token
- Cost model comparing on-premise inference against per-token API billing at your volume
- Operations runbook and handover training for internal IT
Outcomes
What changes
- Confidential data that never leaves your network
- Inference cost fixed at the hardware rather than variable per token
- Known throughput limits, so capacity planning is arithmetic
- No vendor able to change pricing, terms or model behaviour underneath you
- Internal staff able to operate, restart and update the stack without a consultant
Engagement
How we work together
On-premise feasibility and sizing assessment
Deployment and benchmarking project
Hardware procurement advisory
Ongoing model operations retainer
Questions
Straight answers
Is a local model as good as the big hosted ones?
For open-ended reasoning, no, and anyone claiming otherwise is selling. For the tasks most firms actually want, extraction, classification, summarising, drafting from a template, answering from internal documents, current open-weight models are good enough, and the gap narrows every few months. The honest answer comes from testing your task, not from benchmarks.
What hardware do we need?
It depends on model size, quantization level and how many people use it at once. A single workstation GPU with 24GB of VRAM runs a quantized mid-sized model comfortably for a small team; heavier concurrency needs more. Sizing gets settled in the assessment with measured numbers, before anything is purchased.
How does the cost compare to paying per token?
Hosted APIs are cheaper until usage becomes steady and predictable, then the curve crosses. On-premise is a capital cost plus power, so heavy repetitive workloads amortise quickly and light occasional use does not. I build that comparison at your projected volume so the decision is arithmetic rather than preference.
Can it run with no internet connection at all?
Yes. Once the models and dependencies are in place, inference runs entirely offline. That is why this approach suits air-gapped networks, regulated environments and any firm with contractual restrictions on where client data may be processed.
Related
Services that usually go with this
AI Implementation & Automation Consulting
Self-hosted models, real workflows and measurable hours saved. No hype, no per-token surprise bills.
Business Workflow Automation
Find the work being done by hand, count the hours, automate it, then count the hours again.
Custom AI Agents & Assistants
An agent earns its place when it completes a task end to end and hands a person something they can check.
Hiring instead of contracting? See the full-time roles, or the ATS-formatted resume.
Contact
Tell me what you are hiring for
- Call: +1 519 278 5085
- Email: tim@timarmstrong.ca
- Based: London & Southwestern Ontario, Ontario, America/Toronto (Eastern Time)
- Availability: Available nationally across Canada, on-site, hybrid or remote
- Serving: London · Stratford · Woodstock · Kitchener: Waterloo · Cambridge · Guelph · Brantford · Sarnia · Chatham-Kent · Windsor · Toronto · Mississauga and Canada-wide