Skip to main content

Self-hosted · your data never leaves your infrastructure

Fine-tune your own language models — without sending data anywhere

RantAI LLMOps takes a model from dataset to live endpoint on hardware you control. Fine-tune through the UI, evaluate against your own test set, then serve it over an OpenAI-compatible API.

Runs where you host it
Your GPU
Hugging Face or your registry
Any open model
Serves over a standard API
OpenAI-compatible
demo.llmops.rantai.dev/finetuneadmin

learn-4b-full

aisingapore/Gemma-SEA-LION-v4-4B-VL · s3://buku-korpus/learn/v5/

RUNNING
step 130/200
Training loss0.4969
GPU 0 · NVIDIA GB1072°C · 25W
GPU util97%
VRAM66.4 / 121.6 GB
method
SFT · LoRA
rank
16
lr
2e-4

17:10:56 step 126/200 loss 0.5430 lr 2.00e-4 grad_norm 0.499 1.42 it/s

17:10:56 step 127/200 loss 0.5396 lr 2.00e-4 grad_norm 0.692 1.42 it/s

17:10:57 step 128/200 loss 0.4786 lr 2.00e-4 grad_norm 0.616 1.42 it/s

17:10:58 step 129/200 loss 0.4621 lr 2.00e-4 grad_norm 0.458 1.42 it/s

17:10:59 step 130/200 loss 0.4969 lr 2.00e-4 grad_norm 0.679 1.42 it/s

Platform

Everything between a raw model and a working endpoint

Most teams stitch this together from notebooks, shell scripts, and a serving container nobody wants to touch. This replaces that seam.

your-infrastructure
training datamodel weightsinference traffic

outbound

0 B

hosted API

Your data stays put

Training data, model weights, and inference traffic never leave your servers. For universities, government, and regulated industries, that is not a feature — it is the requirement that rules out every hosted API.

asklearnpractice
baseGemma-SEA-LION-v4-4B
GPU 0NVIDIA GB10

Many behaviours, one GPU

LoRA adapters attach to a single frozen base model, so several fine-tuned behaviours can share one card instead of each needing its own deployment.

rantai-ask-4b · grounded98.0%
Qwen2.5-3B · arc_easy76.9%
Qwen2.5-3B · winogrande74.8%

Evaluate before you ship

Run a trained adapter against a held-out set and compare it with the base model, side by side. Promote a version because the numbers moved, not because the loss curve looked pleasant.

vLLMhttp://vllm:8000/v1
Ollamahttp://localhost:11434/v1
llama.cppGGUF export

Bring your own engine

Export to vLLM, Ollama, or llama.cpp and serve over an OpenAI-compatible API. Existing clients point at a new base URL and keep working.

runranklrdataset
practice-4b162e-4learn/v4
learn-4b-full162e-4learn/v5

Every run is reproducible

Hyperparameters, dataset version, and base model are recorded per job. Compare runs, trace a regression back to what changed, and rerun it.

Dataset to endpoint, one place

Import a model and a dataset, fine-tune, evaluate, export, and serve — without switching tools or writing glue scripts between five of them.

import
fine-tune
evaluate
export
serve

Workflow

Four steps, no context switching

The same path every time, whether it is your first adapter or your fortieth.

01

Import

Pull a base model from Hugging Face or your own registry, and point at a dataset in object storage. On-premise corpora never have to be published anywhere public.

ModelsDatasets

Qwen-SEA-LION-v4-4B-VL

Q8_0 · GGUF

Pull

Qwen-SEA-LION-v4-8B-VL

Q8_0 · GGUF

Pull

Gemma-SEA-LION-v4-4B-VL

Downloaded

Dataset

s3://buku-korpus/learn/v5/

02

Fine-tune

Pick LoRA rank, learning rate, epochs, and sequence length, then launch. Loss curves and logs stream into the run view while it trains.

03

Evaluate

Score the result against a held-out set and compare runs head to head, so a version ships on evidence rather than on a hunch.

04

Serve

Export to your engine of choice and expose an OpenAI-compatible endpoint. Adapters are pinned by name, so rolling out a new version changes no client code.

Serving

Ships as an API your code already knows

A fine-tuned model is only useful once something can call it. Models served here speak the OpenAI protocol, so integrating means changing a base URL — not rewriting a client.

  • 01Drop-in OpenAI-compatible API — change the base URL, keep your client
  • 02Adapters selected per request through the model field
  • 03API keys and a model allowlist in front of the engine
  • 04vLLM, Ollama, and llama.cpp as serving targets
import OpenAI from "openai"

const client = new OpenAI({
  baseURL: "https://llm.your-company.internal/v1",
  apiKey: process.env.LLMOPS_API_KEY,
})

const res = await client.chat.completions.create({
  model: "support-agent",   // your fine-tuned adapter
  messages: [{ role: "user", content: "How do I reset my password?" }],
})

response · model: support-agent · engine: vllm

FAQ

Questions teams ask first

Nowhere. The platform runs inside your own environment — training data, model weights, and inference traffic all stay on your hardware. There is no call home, and no vendor-side copy of your corpus.

See it running on your own models

Explore the platform with sample data to see how a run is configured, evaluated, and served. When you are ready to try it on your own models, talk to us.

$ curl https://llm.your-company.internal/v1/models

contact@rantai.dev