Fine-tuning
Fine-tuning starts from a trained model and adapts it to your data. It needs much less compute than training from scratch, which makes short, flexible rentals attractive.
What matters most
Fits in memory
Choose GPUs with enough VRAM for the model and optimiser state, or use memory-saving methods that trade a little speed for space.
Short, bursty jobs
Runs often last hours to days, so per-hour pricing and quick start-up matter more than long contracts.
Simple networking
One node is often enough, so the extra cost of a high-end cluster fabric is usually unnecessary.
Repeatability
You will re-run experiments, so predictable pricing and easy access are valuable.
A typical setup
One node with several GPUs, or even a single large-memory GPU, depending on model size and method.
How it is usually bought
On-demand or short-term capacity is the usual choice. Spot can work if your job checkpoints and can resume.
Common mistakes
- Paying for cluster networking a single-node job doesn’t use.
- Skipping checkpoints on interruptible capacity.
- Underestimating dataset transfer and storage costs.
Terms to know
This is general guidance, not a recommendation for a specific provider or price. Check current offerings and confirm availability and terms with the provider.
Ready to look at real options?
See tracked offerings with their sources and dates, or ask Kova to size it for you.
