Model Serving is the process of making a trained machine learning model available for making predictions on new data. It connects trained models with applications, APIs, or other systems to deliver predictions in real time or in batches.
- Deploys trained models for real-world predictions.
- Supports real-time and batch inference.
- Handles model requests, responses, and resource management.
- Commonly used with tools and platforms such as TensorFlow Serving, TorchServe, and cloud services.

Working
Model serving provides a way for applications to communicate with a trained machine learning model. The application sends new input data to the serving system, which processes the request and returns the model's prediction.
- Application sends new input data
- Serving system receives the prediction request
- Inference code processes the input
- Trained ML model generates a prediction
- Serving system returns the prediction
- Application receives the prediction response
Serving Model using AWS Sagemaker
Step 1: Train the Machine Learning Model
For this demonstration, we will use a very small Linear Regression model.
Create train.py:
from sklearn.linear_model import LinearRegression
import joblib
#training data
X = [[1], [2], [3], [4], [5]]
y = [2, 4, 6, 8, 10]
# Create and train the model
model = LinearRegression()
model.fit(X, y)
# Save the trained model
joblib.dump(model, "model.pkl")
After running the model a new file will appear with the name model.plk, This file is the trained model artifact
Step 2: Test the Model Locally
Before deploying anything to AWS, we should verify that the model works.
Create a simple test:
import joblib
model = joblib.load("model.pkl")
prediction = model.predict([[6]])
print("Prediction:", prediction[0]).
Step 3: Create the Inference Code
Training and inference should be separated. During inference, we do not train the model again. The inference application should:
- Load the trained model.
- Receive input.
- Perform prediction.
- Return the prediction.
For SageMaker's custom inference container, we will expose the standard /ping and /invocations HTTP paths.
Create inference.py:
from flask import Flask, request, jsonify
import joblib
app = Flask(__name__)
model = joblib.load("model.pkl")
@app.route("/", methods=["GET"])
def index():
return jsonify({
"status": "ok",
"endpoints": ["/ping", "/invocations", "/score", "/predict"]
}), 200
@app.route("/ping", methods=["GET"])
def ping():
return "OK", 200
@app.route("/health", methods=["GET"])
def health():
return "OK", 200
@app.route("/invocations", methods=["POST", "GET"])
def invocations():
if request.method == "GET":
return jsonify({"message": "Use POST to send a value."}), 200
data = request.get_json(silent=True)
if data is None:
data = request.form.to_dict()
if not isinstance(data, dict) or "value" not in data:
return jsonify({"error": "Request body must include a JSON field named 'value'"}), 400
try:
value = float(data["value"])
except (TypeError, ValueError):
return jsonify({"error": "Value must be a number"}), 400
prediction = model.predict([[value]])
return jsonify({
"prediction": float(prediction[0])
}), 200
@app.route("/score", methods=["POST", "GET"])
@app.route("/predict", methods=["POST", "GET"])
def aliases():
return invocations()
if __name__ == "__main__":
app.run(host="0.0.0.0", port=8080)
Step 4: Create requirements.txt
Create:
Flaskscikit-learnjoblib
The container will install these dependencies when it is built.
Step 5: Create the Dockerfile
Now we package the model and inference application into a Docker image.
Create Dockerfile:
FROM python:3.11-slimWORKDIR /appCOPY requirements.txt .RUN pip install --no-cache-dir -r requirements.txtCOPY model.pkl .COPY inference.py .EXPOSE 8080CMD ["python", "inference.py"]
Step 6 : Build the Docker Image
Run:
docker build -t model-serving .
Check the image:
docker imagesYou should see:
model-serving
Step 7: Run the Container Locally
Run:
docker run -p 8080:8080 model-servingThe application is now running inside Docker.
Test the health endpoint:
curl http://localhost:8080/pingExpected response:

Step 8: Create an Amazon ECR Repository
- Sign in to the AWS Management Console.
- Search for Amazon ECR.
- Open Elastic Container Registry.
- Select Repositories → Create repository.
- Choose Private.
- Enter a repository name, for example:
model-serving 
Step 9: Push Your Docker Image to ECR
Open the repository you just created.

Click View push commands. All the commands need for pushing an image to ECR will be showed, Following those commands you will be able to push your image to ECR.

Step 10: Create an IAM Role for SageMaker
Go to AWS console and click on IAM, Then click on role and create a role.

Select Trusted entity type: AWS service and for the service select SageMaker.

Copy the Role ARN after creating, because you'll need it when creating the SageMaker model.

Step 11: Create the SageMaker Model
Go to AWS console and search for Amazon SageMaker AI then go to Deployable models

Enter a model name
model-serving-demoFor the IAM role, select the role you created earlier. For the container image, select/use the image URI from your ECR repository.
It will look similar to:
123456789012.dkr.ecr.ap-south-1.amazonaws.com/model-serving:latestThen create the model.

After successfully creating model form the image in ECR you will see something like this.

Step 12: Create an Endpoint
Now go to SageMaker and navigate to Inference and choose Endpoint Give it a name:
model-serving-demo-endpointCreate a new endpoint configuration and select your model. If the goal is to minimize cost, use Serverless Inference if it is available for the model/deployment configuration you're using. Use a small memory allocation appropriate for your model.

Step 13: Deploy the Endpoint
Create the endpoint.

Initially, its status will be something like:
Creating
Wait until the status becomes:
InService
Now your model is deployed.
Cost-Saving Practices
When experimenting with AWS model serving, follow these practices to control costs:
- Use a small ML model that requires minimal compute resources.
- Prefer Serverless Inference for intermittent testing workloads.
- Delete inference endpoints when they are no longer required.
- Remove unused ECR images to reduce storage costs.
- Use only the AWS services required for the model-serving workflow.
- Monitor your AWS Free Tier usage and available credits.
- Set up AWS billing alerts to track unexpected charges.