Model Deployment

Model deployment is the one-time act of putting a trained model into production so it can serve requests, while inference is the ongoing process of that model generating outputs in response to request