Your own AI model server, in place of a bill for every request.
You pay a model vendor for every request your product makes. We set up an open language model on your own machines, give it the request format your code already uses, and test it on your real work before you switch.
A model you run, tested on your own work.
A model chosen by test
A short list of open models, tried on examples from your product.
Both licences read
The serving software and the model files have separate terms. We check both.
Hardware sized first
GPU memory and machine count worked out before any purchase.
The same request format
Where the project supports it, your code changes one address.
Measured under load
Answer times at your real number of requests at the same moment, on your hardware.
A second route
The paid service kept for chosen requests, when that suits your case.
A product that calls a paid model service pays for every request. That is a simple way to start. As use grows, the monthly bill grows with it, and every request sends your customers' text to another company.
The open projects that serve a model from your own machine are listed on our alternatives to the OpenAI API page, among them vLLM, Ollama and llama.cpp. By their own documentation, vLLM and llama.cpp answer in the OpenAI request format, so code written for the paid service can call them after the address is changed, for the request types each project supports.
What we do
Choose the model for your work. A model that summarises support tickets and a model that writes code are different choices. We test a short list on examples from your product and show you the results.
Check the licence of the model files. The serving software and the model are published separately, under separate terms. We read both for your use.
Size the hardware. The model has to fit in the memory of your GPU, and the number of people using it at the same moment decides how many machines you need. We work this out before you buy.
Set up the server and connect your product. With the request format your code already uses where the project supports it, and with connecting code where it does not.
Test under your load. We send your real number of requests at the same moment and measure how long each answer takes. You see the figures for your own hardware.
Keep the paid service as a second route, if you want one. A product can send most requests to its own server and a chosen few to a paid model. We build that choice into the product when it suits your case.
What to check before you decide
Our alternatives page lists where the paid service still wins. The side-by-side test on your own work shows whether any of that matters for you. The parent page, replace a paid AI service, covers when to keep paying.
Common questions
Can we run a language model on our own servers?
Yes. Open projects such as vLLM, Ollama and llama.cpp serve a language model from a machine you control, and our directory lists them as alternatives to the OpenAI API with each project's licence. Reveneau chooses a model for your work, checks that it fits your hardware, sets up the server and connects your product to it.
How much GPU memory does a language model need?
It depends on the model. Each model's own page states its size, and the model has to fit in the memory of the GPU that runs it. The llama.cpp project describes storing a model at lower precision, from 1.5 to 8 bits, to cut the memory it needs. Reveneau works out the memory for your chosen model before you buy hardware.
Will code written for the OpenAI API work with our own model server?
For the request types the project supports, yes. vLLM and llama.cpp both describe an OpenAI-compatible server in their own documentation, so code written for the paid service calls them after the address is changed. Reveneau checks every request type your product uses and writes connecting code for any that the project does not cover.
How many users can one model server handle at the same time?
That depends on the model, the GPU and how long each request is, so the only reliable figure is one measured on your hardware. Reveneau sends your real number of requests at the same moment, measures how long each answer takes, and shows you the results. The test tells you how many machines your busiest hour needs.
Is an open language model as good as a paid one?
It depends on the task, so we give no general answer. Reveneau runs the open model and the paid model on the same examples from your product and gives you both sets of answers to compare. You decide with your own work in front of you, and you can keep the paid service for the requests where it does better.
Can we keep a paid model for some requests and use our own for the rest?
Yes. A product can send most requests to its own model server and a chosen few to a paid model, for example the longest ones. Reveneau builds that rule into the product and shows how many requests go to each side, so you can see the split and change the rule later.
Does our data stay on our machines with our own model server?
With a model server on your own hardware, the text your product sends to the model and the answers it gets back stay on machines you control. Nothing is sent to a model vendor for those requests. If you keep a paid model for some requests, those requests still go to that vendor, and Reveneau marks which ones they are.
Do we need our own engineers to run a model server?
Somebody has to install updates, watch that the server is up, and replace the model when a better one is released. That can be your team, using the instructions and automatic checks Reveneau hands over, or it can be Reveneau under a separate agreement. We agree which before the work starts.
Other services
Replace a paid AI service with one you own.
You pay an AI vendor by the month, by the seat or by use, and the bill grows with your business. We set up an open-source replacement on your own servers, connect it to your product, and build the software around it. At the end the servers, the data and the code are yours.
Your own company chat assistant, in place of a seat for every person.
You pay a monthly seat for every person who uses a chat assistant. We set up an open-source chat assistant on your own servers, with your company sign-in, and connect it to the models you choose.
Your own search over company documents, on servers you control.
You pay for a tool that searches your company's files and answers questions from them. We set up an open-source replacement on your own servers, connect it to where your documents live, and keep each person's view limited to what they are allowed to see.
Your own automation platform, in place of a bill for every task.
You pay an automation service for every task it runs, and the automations your business depends on live in someone else's account. We set up an open-source automation platform on your own servers and rebuild your automations on it.
Your own voice AI: speech, transcripts and phone agents on your servers.
You pay a voice service for every character it speaks or every minute it listens. We set up open-source speech models on your own hardware, connect them to your product or your phone system, and test them in your language before you switch.
Your own AI coding tools, with your code kept on your machines.
You pay a seat for every engineer who uses an AI coding assistant, and parts of your source code are sent to the vendor with each request. We set up open-source coding tools for your team, connected to a model you choose, including one on your own servers.
Build with us
Release software with confidence.
Move faster without lowering your standards. A software development partner whose senior teams integrate into how you decide and stay accountable through delivery.
An MVP development company built for your next milestone.
An MVP development company for founders: senior judgment and focused sprints that strengthen the business, not just release features.
A build partner for the companies you believe in.
Technical due diligence and senior leaders who reduce the technical risks across your portfolio.