Units / EPM5017
EPM5017 · Machine learning for biostatistics
2026 Handbook6 credit pointsLevel 5Department of Epidemiology and Preventive Medicine
Last checked: 23 Aug 2026 UTCOverview
Recent years have brought a rapid growth in the amount and complexity of health data captured. Among others, data collected in imaging, genomic, health registries and personal devices call for new statistical techniques in both predictive and descriptive learning. Machine learning algorithms for classification and prediction complement classical statistical tools in the analysis of these data. This unit will cover modern machine learning methods particularly useful for large and complex data. Topics include, classification trees, random forests, model selection, lasso, bootstrapping, cross-validation, generalised additive modelling, and regression splines. The statistical software R package will be used throughout the unit.
Offerings
| Campus | Teaching period | Mode |
|---|---|---|
| Alfred Hospital | Second semester | Teaching is all online (ONLINE) |
Assessment
The Handbook does not list a final examination among the assessment items. That is not a guarantee there is none.
| # | Assessment | Type | Weight | Hurdle |
|---|---|---|---|---|
| 1 | 2 x Theoretical exercises | Exercise | 80% | Threshold |
| 2 | 2 x Short exercises | Exercise | 20% | — |
Assessment in this unit includes hurdle assessment tasks. Failure of any hurdle assessment task may result in failure of the unit
Assessment details may change. Please refer to the assessment information in Moodle closer to the start of the teaching period.
Requisites
Learning outcomes
- Describe situations where machine learning methods can offer advantages over traditional statistical modelling approaches to data analyses in health applications
- Recognise and explain the differences between the goals of description and prediction
- Determine and implement appropriate machine learning approaches for description and prediction in real-world health applications
- Measure and explain the uncertainty of the results of analyses using machine learning approaches
- Interpret the results of analyses using machine learning in light of the assumptions required, the quality of input data, and the sensitivity to the specific technique implemented
- Critically appraise current literature concerning machine learning applications for classification or prediction in health
- Effectively communicate in language suitable for the scientific community the results of analyses using machine learning methods
Workload
Off campus: Twelve hours per week, consisting of (on average) 4 hours per week for reading core material, 4 hours per week completing exercises (manual, computer-based, or on-line), 2 hours per week for on-line communication via discussions, and 2 hours per week for assignment preparation. No residential component is required for this unit.
Ask about EPM5017
Answered from the Handbook fields above — no AI, no guessing. Every answer links back to the source.
Community discussions about EPM5017
CommunityStudent experience, not official rules. Nothing here changes what the Handbook says.