Product Management

7 Critical Machine Intelligence Exams and The Hidden Link of MLOps with Product Management

It is time machines take exams before being unleashed on the world! A great machine learning product becomes extraordinary through testing...

Jean Voigt
December 10, 20218 min read
Photo by Akshay Chauhan on Unsplash
Photo by Akshay Chauhan on Unsplash

3:15 pm the day before Thanksgiving. You are cleaning up your inbox for the long weekend. Then, your sales department head calls: _"What the h*** happened to sales lead prioritizations???"_No recent deployment, and the data checks out as well. It looks like your Thanksgiving plans suddenly changed!

This article will examine the link between product management and MLOPs in a fast-paced rundown of key organizational, technical, and, most importantly, exams machine learning models should pass.

The past 12 months have seen many machine learning operations tools gaining prominent popularity. Interestingly, one feature is notably absent or rarely mentioned in the discussion: quality assurance.

Academia has already initiated research in machine learning system testing. In addition, several vendors provide data quality support or leverage data testing libraries or data quality frameworks. Automated deployment does exist as well in many tools. But how about canary deployments of models and whatever happened with unit and integration testing in the machine learning universe?

Many of these quality assurance proposals originate from an engineering mindset. However, more and more specialists without an engineering background perform a lot of model engineering. Further, recall that a separate person or team frequently runs the quality assurance activities. Supposedly so that engineers can place their trust in others to catch mistakes. More cynical characters might insist that engineers need to be controlled and checked.

Before getting any deeper into machine intelligence exams, allow for a quick discussion of an engineering mindset and the perils of quality assurance delegation.

Science is not engineering... that is OK!

Via Imgur
Via Imgur

Most of the time, in science, it is totally fine if things work once, in a controlled environment, and often, experiments may actually fail. In engineering, there is no such luxury. For all intents and purposes, many systems run forever, giving rise to quality assurance. This mindset matters. Data scientists experiment and rarely need reproducible results outside of regulatory or audit contexts. Engineers have to provide reproducible results reliably over extended periods. That doesn't make science any easier than engineering, but it sheds light on the difference in objectives. Therefore, shifting the mindset from a data science perspective to an engineering perspective provides much-needed clarity. Next, let's discuss who might perform the quality assurance work.

Test, test, test.... does it work yet?

Photo by @felipepelaquim on Unsplash
Photo by @felipepelaquim on Unsplash

With that engineering perspective in mind, let's examine the often observed separation of quality assurance from engineering. Quality assurance comes in many forms and shapes but boils down to testing. Unless you really worry about guard rail programming, test automation is the way to go; Patric's article explains why. Clearly, before testing anything, you need to know what to expect. For machine learning that's typically operationalized into an objective function - in laymen's terms: Figure out how success looks like.

Two essential organizational aspects dominate testing:

Delegation cannot be wrong, right? So why not delegate and have a specialist test the model? There are well-argued points to separate test engineering from development engineering for the sake of specialized test automation orchestration tasks. However, a culture of distrust resulting from dividing engineers and testers is essentially unhealthy. So, ideally, engineers ensure to test their own model, test engineers look to automate the testing effort. After all, engineers know best how to properly check inbound data, so that model assumptions don't get corrupted. Some rotation between roles might help keep everyone well informed on the latest practices. Note that the rotation is between roles on the same team. Rotation between teams is far less effective due to ramp-up/ramp-down efforts.

Whether to use formal or data-centric model testing will often depend on whether formal validation is possible and economical. For most situations, data-centric testing would work very well.

Let's recap: To build reliable machine learning models for production operations, data-centric automated testing is an excellent solution in most situations.

Luckily, many fantastic software engineering tools in the CI/CD universe can be repurposed for model testing. Some emerging products may even provide built-in model testing features in due time. If you do not want to wait for that, you can build your own, or, with some help from the DevOps team, repurpose existing tools:

  1. Split and version datasets

  2. Control inbound data quality right into your Grafana

  3. Use pre-commit hooks to ensure data and model tests are there

  4. Leverage preferred CI tool/testing framework to run tests on model & data

  5. Establish model monitoring metrics

  6. Canary deploy models same as other code

The remainder of the article illustrates a few note-worthy aspects to consider when designing effective tests for machine learning systems. If it helps, these aspects are to machine learning models what exams are to school children

Functional testing

Photo by Artem Bryzgalov on Unsplash
Photo by Artem Bryzgalov on Unsplash

Functional unit and integration tests are often indispensable. Especially when breaking down a larger model into smaller sub-models dedicated to a specific task, say an image and a text classification model as part of a larger article rating model. Functional tests should ensure choosing the proper computation methods and selecting the correct parameters and data for each sub-model. The functional testing might also include the input data quality testing, including testing for model metric sensitivity to structurally sound but statistically degenerated data.

Performance testing

Photo by Mikail McVerry on Unsplash
Photo by Mikail McVerry on Unsplash

Testing for performance is the most standard test scenario for machine learning. In a way, this is a pre-condition of all other tests looking for sensitivity to performance metrics. However, specific metrics may be important to a model design or threshold limits for a particular training cohort, test, or validation datasets. Hence, considering such choices in the test case design improves testing quality.

Label quality sensitivity testing

Photo by Pop & Zebra on Unsplash
Photo by Pop & Zebra on Unsplash

There are six types of label data issues that are worthwhile to test. When testing for label data sensitivity of the model, the label data can either be curated manually, or tools suck as Snokel can generate synthetic training data for the model test cases. In either event, tests are looking for crucial model metric sensitivity to each class of label data issue. Due to the complexity of label data management and potential synthetic data generation, label quality testing is probably the more complex set of testing activities.

Ethical and regulatory testing

Photo by Tingey Injury Law Firm on Unsplash
Photo by Tingey Injury Law Firm on Unsplash

Designing test cases to minimize unintended consequences, including regulatory rule violation or reducing ethical concerns, is becoming more and more critical. The regulatory domain applicable for your model may vary. However, independent of that, the sensitivity to purposefully biased training/test data will help determine how and where one version of the model introduces more or less bias.

Consistency testing

Photo by Bernard Hermant on Unsplash
Photo by Bernard Hermant on Unsplash

To borrow an analogy, a nail transformed by a hammer still remains a nail. The same should be valid for data flowing through a model pipeline. Yet, data invariants more often than not get violated. Software engineers test for these.

With machine learning, there is a second consistency aspect. The function of the nail needs to remain consistent. If a factor or parameter of a model resulted in a specific (measurable) behavior in the last model version, the behavior should be the same, except for intentional changes. Indeed, the source for how factors change is somewhat complex. In addition, there is a significant impact from interactions and data used for training and testing. However, there are likely core believes of your model you don't want to change for good reasons. For example, in most cases, you don't want to bend the law of physics, and higher objects should result in higher volume too.

Hyperparameter corner cases

Photo by Volodymyr Hryshchenko on Unsplash
Photo by Volodymyr Hryshchenko on Unsplash

It may be illustrative to understand where the model configuration sits in the hyperparameter space, especially when using automated hyperparameter tuning methods. Verifying model results for corner cases is one option to provide clarity on hyperparameter choices. If a model parameter configuration is very close to a specific corner case, perhaps a manual re-evaluation of parameter settings might be appropriate. The tuning algorithm configuration might be too aggressive, particularly when the tuning algorithm settings are determined automatically.

Drift tests

Photo by Ralfs Blumbergs on Unsplash
Photo by Ralfs Blumbergs on Unsplash

As the world evolves, so do machine learning models, or rather they should. While model monitoring may detect a drift in a production system, more is possible when deploying a new model version. Running the new model against multiple time-sliced datasets to estimate the drift inherent in the new model version indicates future drift expectations. Introducing more significant drift would impact the scheduled model re-train and update frequency and is an essential input to model operations.

Summary

Establishing sound MLOps is already a complicated and complex undertaking; adding testing does not make it any easier. As for any machine learning effort, the critical aspect is to define a clear and measurable objective, including model metrics to quantify if the objective is reached and to what degree.

The rewards are often worth the invested effort, not only because it helps to prevent disruptive incident management efforts. The data collected as part of testing efforts allows for a more controlled model design.

However, the test data are dual-use in nature. They can be vital for audit and regulatory discussions and help with marketing communication to demonstrate specific product quality attributes. Such results are far more illustrative than industry benchmarks as a communication device. Especially when the machine learning product targets a technical buying center, such data points are great conversation starters in presentations or social media content. Inviting potential clients to examine testing strategies allows creating confidence in the product. Without rigorous testing, this opportunity is overlooked entirely.

Exams for machine intelligence might well save a Thanksgiving gathering and potentially improve the marketing of your product. Communicating much more objectively about a product using test data results provides an essential link between MLOps and product management.

Further reading

  1. Testing machine learning based systems: a systematic mapping

  2. Test-Driven Machine Learning

  3. Effective test for machine learning systems

  4. Efemarai machine learning model testing platform

  5. A Practical Guide to Maintaining Machine Learning in Production

Related Articles