Model Serving

Last Updated : 21 Sep, 2026

Model Serving is the process of making a trained machine learning model available for making predictions on new data. It connects trained models with applications, APIs, or other systems to deliver predictions in real time or in batches.

  • Deploys trained models for real-world predictions.
  • Supports real-time and batch inference.
  • Handles model requests, responses, and resource management.
  • Commonly used with tools and platforms such as TensorFlow Serving, TorchServe, and cloud services.
frame_3902

Working

Model serving provides a way for applications to communicate with a trained machine learning model. The application sends new input data to the serving system, which processes the request and returns the model's prediction.

  • Application sends new input data
  • Serving system receives the prediction request
  • Inference code processes the input
  • Trained ML model generates a prediction
  • Serving system returns the prediction
  • Application receives the prediction response

Serving Model using AWS Sagemaker

Step 1: Train the Machine Learning Model

For this demonstration, we will use a very small Linear Regression model.

Create train.py:

Python
from sklearn.linear_model import LinearRegression
import joblib

#training data
X = [[1], [2], [3], [4], [5]] 
y = [2, 4, 6, 8, 10]

# Create and train the model 
model = LinearRegression() 
model.fit(X, y)

# Save the trained model 
joblib.dump(model, "model.pkl")

After running the model a new file will appear with the name model.plk, This file is the trained model artifact

Step 2: Test the Model Locally

Before deploying anything to AWS, we should verify that the model works.

Create a simple test:

Python
import joblib

model = joblib.load("model.pkl")

prediction = model.predict([[6]])

print("Prediction:", prediction[0]).

Step 3: Create the Inference Code

Training and inference should be separated. During inference, we do not train the model again. The inference application should:

  1. Load the trained model.
  2. Receive input.
  3. Perform prediction.
  4. Return the prediction.

For SageMaker's custom inference container, we will expose the standard /ping and /invocations HTTP paths.

Create inference.py:

Python
from flask import Flask, request, jsonify
import joblib

app = Flask(__name__)

model = joblib.load("model.pkl")


@app.route("/", methods=["GET"])
def index():
    return jsonify({
        "status": "ok",
        "endpoints": ["/ping", "/invocations", "/score", "/predict"]
    }), 200


@app.route("/ping", methods=["GET"])
def ping():
    return "OK", 200


@app.route("/health", methods=["GET"])
def health():
    return "OK", 200


@app.route("/invocations", methods=["POST", "GET"])
def invocations():
    if request.method == "GET":
        return jsonify({"message": "Use POST to send a value."}), 200

    data = request.get_json(silent=True)
    if data is None:
        data = request.form.to_dict()

    if not isinstance(data, dict) or "value" not in data:
        return jsonify({"error": "Request body must include a JSON field named 'value'"}), 400

    try:
        value = float(data["value"])
    except (TypeError, ValueError):
        return jsonify({"error": "Value must be a number"}), 400

    prediction = model.predict([[value]])

    return jsonify({
        "prediction": float(prediction[0])
    }), 200


@app.route("/score", methods=["POST", "GET"])
@app.route("/predict", methods=["POST", "GET"])
def aliases():
    return invocations()


if __name__ == "__main__":
    app.run(host="0.0.0.0", port=8080)

Step 4: Create requirements.txt

Create:

Flask
scikit-learn
joblib

The container will install these dependencies when it is built.

Step 5: Create the Dockerfile

Now we package the model and inference application into a Docker image.

Create Dockerfile:

FROM python:3.11-slim

WORKDIR /app

COPY requirements.txt .

RUN pip install --no-cache-dir -r requirements.txt

COPY model.pkl .
COPY inference.py .

EXPOSE 8080

CMD ["python", "inference.py"]

Step 6 : Build the Docker Image

Run:

docker build -t model-serving .
Screenshot-2026-09-21-152027

Check the image:

docker images

You should see:

model-serving
Screenshot-2026-09-21-152135

Step 7: Run the Container Locally

Run:

docker run -p 8080:8080 model-serving

The application is now running inside Docker.

Test the health endpoint:

curl http://localhost:8080/ping

Expected response:

Screenshot-2026-09-21-153040

Step 8: Create an Amazon ECR Repository

  1. Sign in to the AWS Management Console.
  2. Search for Amazon ECR.
  3. Open Elastic Container Registry.
  4. Select Repositories → Create repository.
  5. Choose Private.
  6. Enter a repository name, for example:
model-serving 
Screenshot-2026-09-21-153633

Step 9: Push Your Docker Image to ECR

Open the repository you just created.

Screenshot-2026-09-21-154152

Click View push commands. All the commands need for pushing an image to ECR will be showed, Following those commands you will be able to push your image to ECR.

Screenshot-2026-09-21-154506

Step 10: Create an IAM Role for SageMaker

Go to AWS console and click on IAM, Then click on role and create a role.

Screenshot-2026-09-21-155256

Select Trusted entity type: AWS service and for the service select SageMaker.

Screenshot-2026-09-21-155346

Copy the Role ARN after creating, because you'll need it when creating the SageMaker model.

Screenshot-2026-09-21-155644

Step 11: Create the SageMaker Model

Go to AWS console and search for Amazon SageMaker AI then go to Deployable models

Screenshot-2026-09-21-160631

Enter a model name

model-serving-demo

For the IAM role, select the role you created earlier. For the container image, select/use the image URI from your ECR repository.

It will look similar to:

123456789012.dkr.ecr.ap-south-1.amazonaws.com/model-serving:latest

Then create the model.

Screenshot-2026-09-21-161135

After successfully creating model form the image in ECR you will see something like this.

Screenshot-2026-09-21-172619

Step 12: Create an Endpoint

Now go to SageMaker and navigate to Inference and choose Endpoint Give it a name:

model-serving-demo-endpoint

Create a new endpoint configuration and select your model. If the goal is to minimize cost, use Serverless Inference if it is available for the model/deployment configuration you're using. Use a small memory allocation appropriate for your model.

Screenshot-2026-09-21-163651

Step 13: Deploy the Endpoint

Create the endpoint.

Screenshot-2026-09-21-172938

Initially, its status will be something like:

Creating
Screenshot-2026-09-21-174435

Wait until the status becomes:

InService

Now your model is deployed.

Cost-Saving Practices

When experimenting with AWS model serving, follow these practices to control costs:

  • Use a small ML model that requires minimal compute resources.
  • Prefer Serverless Inference for intermittent testing workloads.
  • Delete inference endpoints when they are no longer required.
  • Remove unused ECR images to reduce storage costs.
  • Use only the AWS services required for the model-serving workflow.
  • Monitor your AWS Free Tier usage and available credits.
  • Set up AWS billing alerts to track unexpected charges.
Comment