Implementing Good Machine Learning Practices for AI Devices

Implementing Good Machine Learning Practices for AI Devices
07-Jul-2025

Good Machine Learning Practice (GMLP) for AI Medical Devices

Implementing Good Machine Learning Practices for AI Devices

Artificial Intelligence (AI) and Machine Learning (ML) are rapidly transforming healthcare. AI/ML models are making significant contributions to faster diagnosis, personalized treatments, and efficient care. But how do we make sure that AI used in medical devices is reliable, fair, and safe?

To answer this question, the FDA, Health Canada, and the UK MHRA jointly published 10 Guiding Principles for Good Machine Learning Practice (GMLP). These principles help ensure AI/ML based medical devices are developed responsibly across their full lifecycle.

The 10 principles are:

1. Leverage Multidisciplinary Expertise Throughout the Product Lifecycle:

AI in healthcare isn't a one person job. Developing AI in healthcare requires input from clinical experts, data scientists, software engineers, quality and regulatory professionals, and user experience designers. Each stakeholder plays a critical role in ensuring that the final product is clinically relevant, technically robust, and operationally feasible.

ECG example: Cardiologists define what patterns indicate arrhythmia, engineers translate that into algorithms, and UI designers ensure it's easy for clinicians to use.

2. Implement Good Software Engineering and Security Practices:

Clean, well documented code, reliable infrastructure, and strong cybersecurity are the foundation of safe AI/ML healthcare products. These practices involve a clear and careful design and risk management process. Engineers must maintain clean architecture, version control, quality assurance, and strong cybersecurity for developing a safe, reliable, and secure product.

ECG example: The software must be protected from unauthorized access and handle real time ECG data without crashing or introducing errors.

3. Ensure Clinical Study Participants and Data Sets are Representative of the Intended Patient Population:

Clinical study participants and data sets should reflect the people the device is meant to help. This means including a wide range of individuals by age, sex, race, and other factors, in a large enough sample. Doing this helps make sure the device works well for everyone, avoids bias, and shows where the model might not perform as expected.

ECG example: If the model is intended for broad use, it should be trained on ECG data from both men and women across various age groups and ethnic backgrounds.

4. Ensure Training Data Sets are Independent of Test Sets:

Training and evaluation data should be fully independent to ensure reliable performance metrics. This reduces the risk of data leakage or misleading results.

ECG example: ECGs from the same individual or recorded at the same clinical site should not appear in both the training and validation sets.

5. Ensure Reference Datasets are Based Upon Best Available Methods:

Model accuracy depends heavily on the quality and reliability of labeled data. Reference standards should follow best available clinical guidelines and, where necessary, be reviewed by expert clinicians.

ECG example: Diagnoses of arrhythmia should be confirmed by board certified cardiologists, rather than relying solely on automated annotations.

6. Tailor the Model Data to the Available Data and Ensure it Reflects the Intended Use of the Device:

The model should match the available data and be built to reduce risks like overfitting and security issues. Its design should reflect how it will be used in real patients and settings, with clear goals to ensure safe and effective performance.

ECG example: If the tool is for wearable ECG monitors, it must tolerate noise and work on shorter recordings.

7. Focus Should be Placed on the Performance Human AI Team:

The focus should be on supporting clinical decision making. Interfaces must be designed for clarity, interpretability, and context specific use.

ECG example: Rather than issuing a binary result, the AI should highlight affected signal segments and display confidence scores to support clinical interpretation.

8. Conduct Testing Under Clinically Relevant Conditions:

AI that works in a lab might fail in the real world. The model must be tested in settings that mirror real world clinical environments, including diverse user groups, typical workflows, and potential confounding factors.

ECG example: Test it in hospitals, primary care clinics, even noisy home environments, wherever it's going to be deployed.

9. Provide Users with Clear/Essential Information:

Users should have easy access to clear and relevant information, like how the product is meant to be used, how well it works for different groups, what data it was trained on, and any limitations. They should also be informed about updates, how to use the interface, and how it fits into their workflow. A way to share feedback or concerns should also be available.

ECG example: The interface should disclose what types of arrhythmias can be detected, known limitations, and how to act on the AI's output.

10. Monitor Model Performance Post Deployment and Manage Re Training Risks:

Post market monitoring is essential to ensure ongoing safety and effectiveness. If models are periodically updated or retrained, there must be controls to manage data drift, unintended bias, and performance degradation.

ECG example: After deployment, the model's performance should be continuously evaluated on new data from various devices and locations, with a formal process to update or roll back versions as needed.

easyQ Editorial Team

easyQ Editorial Team

Provides expert insights on medical device quality management, regulatory compliance, and eQMS solutions to help MedTech companies simplify compliance and improve quality processes.

Start Your Smart Compliance Journey

Get expert guidance and simplify your compliance process today — talk to our team about how easyQ fits your QMS.

Talk to Our Experts
easyQ compliance experts