Paid
Replicate
Run open models behind an API without provisioning a single GPU
What it is
The practical way to use image, audio and video models that the text-first providers do not serve, plus fine-tuning on your own data. Billing is per second of compute rather than per token, which means a slow model costs more for identical output, and a model nobody has called recently pays a cold-start penalty before it answers.
