AI-powered disease incidence forecasting for hospital planning

Predict disease incidence. Prepare hospital resources before the peak.

Be Healthy transforms clinical, epidemiological and operational data into forecasts, alerts and planning signals for hospital management and clinical teams - helping hospitals move from reactive management to predictive operational planning, alongside an explainable decision support layer for doctors. Built as support for management and clinicians, not a replacement.

Proof-of-concept analytics platform · not a certified medical device

Two complementary layers

One platform, two audiences

Be Healthy is built as a management layer for hospital planning and a clinical layer for doctors - sharing the same explainable AI foundation.

Layer 1 · Hospital Intelligence

Disease incidence forecasting and peak detection that turns epidemiological signals into planning decisions for beds, staff and procurement - built for hospital management and executive teams.

Explore Disease Forecasting

Layer 2 · Doctor Assistant

Patient-level risk assessment with explainable, SHAP-style recommendations - a clinical decision support assistant for doctors, not a replacement for clinical judgment.

Explore Doctor Assistant

The problem hospitals face

Hospitals often react after the pressure has already arrived

Purchases made too late

Supplies, medicines, diagnostics and protective equipment are often ordered under pressure, once demand has already risen.

Planning based on history, not forecasts

Beds, doctors, nurses and diagnostic teams are typically planned from historical reports rather than near-term forecasts.

Seasonal peaks create overload

Influenza, COVID-19, RSV and other seasonal or chronic-disease peaks create operational overload that arrives faster than staffing and supply plans can adjust.

No early-warning dashboard

Managers often lack a single dashboard connecting epidemiological signals to operational resources - moving hospitals from reactive management to predictive planning is the goal.

Hospital Command Center

An early-warning and planning dashboard for management

Disease incidence forecasting, peak detection and demand signals for beds, staff and supplies - brought together in one decision-support dashboard for hospital leadership.

Predicted disease pressure 72/100
Expected peak week Week 6
Bed demand signal Elevated
Staff demand signal Elevated
Procurement readiness 68%
Alert level High

Illustrative demo - figures shown here are not live hospital data. PoC only.

Layer 2 · Clinical decision support

Doctor Assistant: explainable AI at the patient level

Alongside hospital-level forecasting, Be Healthy offers a clinical decision support assistant - patient-level risk assessment with explainable, SHAP-style recommendations. The Heart Disease Risk Assistant is one example of this functionality.

  • Explainable by design
  • Guideline-aligned · ESC / ACC / AHA
  • Support, not replacement

Example Modules

See both layers in action

Illustrative, browser-only PoC demonstrators - no backend, no patient data sent.

What it shows

Browser-only PoC preview of seasonal disease incidence forecasting - no backend, no real patient data.

Why it matters

Turns epidemiological signals into an early planning window before demand peaks.

Who it's for

Hospital management and executive planning teams.

Competition

Positioning in the healthcare solutions market

Traditional HIS/EHR systems

  • No predictions
  • No AI
  • Historical reporting only

Point AI models

  • One problem = one model
  • Lack of integration
  • Lack of explainability
Two-layer platform

.be healthy

  • Hospital-level forecasting plus per-patient decision support
  • Explainable predictions by design
  • Built to extend across disease areas and hospital departments

Key benefits

Real value for patients, physicians and the healthcare system

For care quality

Clinical benefits

  • Earlier identification of high-risk patients
  • Designed to help reduce severe complications through earlier intervention
  • Better-informed treatment decisions, supported by data
  • Greater clinical trust in AI through explainability
For leadership

Strategic benefits

  • Competitive advantage for adopting facilities
  • Improved readiness for AI regulation in medicine
  • Scalable to additional diseases and modules
  • A foundation for broader digital transformation in healthcare
For operations

Operational benefits

  • Supports better procurement timing
  • Helps prepare beds and staff before peak demand
  • Designed to support cost-aware planning - may reduce emergency procurement after validation
  • Faster response to emerging health trends

Compliance notice

Important notice

Be Healthy is a proof-of-concept analytical platform. It is not a certified medical device and is not intended for independent clinical decision-making.

  • The system does not diagnose, treat, cure or prevent any disease.
  • All benchmark figures (accuracy, recall, AUC) are internal PoC / demonstration results, not certified clinical performance.
  • Predictions are intended to support - never replace - the judgment of qualified healthcare professionals.
  • Clinical deployment would require large-scale validation, regulatory review and institutional integration.
PoC only Validation required Not for clinical use

Layer 1 · Hospital Intelligence

Disease incidence forecasting for hospital planning

Be Healthy helps hospitals and healthcare managers anticipate disease incidence peaks and translate forecasts into operational decisions - planning beds, staff and procurement before pressure arrives, not after.

Hospital board Hospital management Operational & medical directors Public health decision-makers

Executive problem

Hospitals often react after disease pressure has already increased

Purchases made too late or under pressure

Supplies, medicines, diagnostics and protective equipment are frequently ordered reactively, once pressure is already visible on the wards.

Beds and staff planned from historical reports

Staffing and capacity decisions are typically based on past reporting cycles rather than near-term forecasts of what's coming.

Seasonal peaks create overload

Peaks such as influenza, COVID-19, RSV and other respiratory or chronic-disease burdens create operational overload for hospitals unprepared for the timing.

No early-warning dashboards

Managers often lack dashboards that connect epidemiological signals with the operational resources needed to respond to them.

Systemic effect

Reactive, siloed management

Without shared planning signals, clinical and operational teams coordinate their response after the fact rather than ahead of it.

Management solution

An early-warning and planning platform

Be Healthy is designed as an early-warning and planning layer that sits alongside hospital operations.

  • Disease incidence forecasting - near-term projections of expected case volumes.
  • Peak detection - early identification of when pressure is expected to peak.
  • Demand signals - for beds, staff and supplies, derived from the forecast.
  • Resource planning dashboard - a single view connecting forecasts to operational decisions.
  • Procurement timing support - planning signals for when to order supplies, medicines and equipment.
  • Scenario comparison - comparing illustrative disease-pressure scenarios side by side.
  • Alerts & recommendations - planning signals surfaced to management as pressure builds.

Hospital Command Center

From forecast to operational decision

The Hospital Command Center brings predicted disease pressure, expected peak week, bed and staff demand signals, procurement readiness and departmental pressure together in one decision-support dashboard - illustrative in this PoC, designed for board-level clarity.

Operational benefits

Moving from reactive management to predictive planning

  • Supports better procurement timing
  • Supports resource planning across beds, staff and supplies
  • Helps prepare beds and staff before peak demand
  • Helps reduce avoidable last-minute decisions
  • Helps management coordinate clinical and operational response
  • Designed to support cost-aware planning - may reduce emergency procurement after validation

See it in action

Explore the forecasting demonstrator

An illustrative, browser-only PoC demonstrator of disease incidence forecasting.

About this module

A high-level look at what Disease Forecasting represents as a planning layer, and the context behind the signals it illustrates.

What this demo illustrates

This demonstrator shows how seasonal disease-incidence signals can be translated into early planning cues for hospital management. It focuses on trend direction, expected pressure and preparation timing rather than certified epidemiological forecasting.

Data & validation context

The page uses an illustrative reference scenario and browser-generated demo signals. It is intended to communicate the planning concept, not to represent live surveillance data or validated public-health forecasts.

Practical applications
  • Early warning before seasonal pressure increases
  • Planning beds, staff and procurement before peak demand
  • Comparing regional or scenario-level pressure patterns
  • Supporting management discussions with transparent signals
Limitations
  • PoC only
  • No live hospital or public-health data
  • Requires external validation before operational deployment
  • Not a certified medical device

Layer 1 · Hospital Intelligence

Hospital Command Center - Illustrative Planning Dashboard

Predicted disease pressure, expected peak week, bed and staff demand signals, procurement readiness and departmental pressure in one decision-support view. Select an illustrative scenario below.

Illustrative demo PoC No backend Requires validation before operational deployment
ModeIllustrative PoC dashboard
DataIllustrative / synthetic scenarios
ScopeHospital-level planning signals
UseRequires validation before operational deployment

Scenario & Planning Inputs

Illustrative input

Scenario: Influenza-like seasonal wave

Advanced parameters

Illustrative demo · synthetic planning signal, not live hospital data.

Disease Pressure Forecast

Synthetic curve
Illustrative disease pressure forecast Synthetic pressure curve rising toward the expected peak week, based on the current planning inputs. Shaded band marks the procurement planning window. Peak
Disease pressure (synthetic) Expected peak week Procurement planning window

Procurement Cost Outlook

Synthetic estimate
Order now $--
Delayed cost $--
Risk premium +--%
Avoidance $--

-

Synthetic estimate · not a financial forecast.

Operational Summary

Moderate
Disease pressure --/100
Peak week --
Bed demand --
Staff demand --
Procurement readiness --%
Recommended management action Moderate

-

Illustrative rule-based demo · not a validated operational or clinical protocol.

Procurement Timing

Cost-aware planning signal
Supply cost pressure --/100
Supplier lead time -- days
Readiness --%
Suggested procurement window Review within 7 days

-

Bed & Staff Readiness

Illustrative capacity view
Bed capacity status -- --
Staff coverage status -- --

Departmental Pressure

Illustrative, all departments - scales with disease pressure
Illustrative demo. Figures shown here are not live hospital data - computed live in your browser from a transparent rule-based heuristic for demonstration purposes only. This is a PoC planning signal, not an operational deployment, and requires validation before real-world use.

About this module

What the full Command Center is designed to become, and what this illustrative PoC does not yet do.

What this illustrates

A scenario preset (baseline, influenza-like, COVID-19-like, RSV-like) loads a starting set of planning inputs - disease pressure, bed occupancy, staff capacity, procurement readiness, supply cost pressure, supplier lead time, expected peak week and regional trend - which you can then fine-tune with the sliders. Every KPI, the forecast chart, the alert level, the recommended management action, the procurement cost estimate, the procurement timing suggestion, the bed/staff readiness cards and the departmental pressure grid all recompute live from those inputs via a transparent rule-based heuristic (not a trained model), entirely in your browser.

Management value
  • Earlier procurement review
  • Reduced emergency purchasing risk
  • Better preparation before peak demand
  • Shared operational view for management teams
Intended data sources (future)
  • Epidemiological surveillance and regional incidence data (see the Disease Forecasting demonstrator)
  • Hospital EHR / admissions data for bed and staff occupancy trends
  • Procurement / ERP systems for supply and equipment readiness, cost and supplier lead times
  • Historical seasonal patterns for peak-week estimation
Limitations
  • All figures come from illustrative presets and slider inputs, not from live or historical hospital data.
  • The alert level, recommendation, procurement timing and estimated procurement costs are a transparent demo heuristic (simple weighted rules), not a validated clinical, operational or financial model.
  • The procurement cost figures (budget, delayed cost, savings) are a synthetic illustrative estimate scaled from a placeholder base budget - not a real quote, invoice or financial forecast.
  • Department pressure values are demonstration numbers, scaled illustratively from the selected scenario - not derived from real occupancy or staffing systems.
  • This dashboard concept requires validation before any operational or clinical deployment.
  • Not a substitute for institutional emergency planning or public health guidance.

Layer 2 · Clinical decision support

Doctor Assistant - explainable AI at the patient level

Alongside hospital-level forecasting, Be Healthy offers a clinical decision support assistant: patient-level risk assessment with explainable, SHAP-style recommendations. Built as AI support for doctors, not a replacement for clinical judgment.

What problems do we solve?

Challenges in clinical decision-making

Predictive gap

Reactive, not predictive

Clinical decisions are often made after symptoms appear, rather than ahead of the curve.

Why it matters: intervening after symptoms appear means missed windows for prevention.

Alerting gap

No early warnings

Chronic conditions and seasonal outbreaks such as influenza too often go unflagged until they escalate.

Why it matters: outbreaks and chronic flare-ups are caught later than they could be.

Trust barrier

Limited interpretability

Opaque AI scores are hard for clinicians to explain - and harder still to trust with a patient in front of them.

Why it matters: clinicians hesitate to act on scores they can't explain to patients.

Compliance burden

Regulatory complexity

Bringing AI into clinical workflows means navigating GDPR, MDR and evolving regulatory expectations.

Why it matters: compliance gaps slow adoption even when the model performs well.

Solution

Be Healthy: a clinical decision support AI platform

Be Healthy is designed as an integrated analytics layer that sits alongside existing clinical systems - combining risk prediction, explanation and visualization in one modular platform.

  • Clinical decision support AI platform - risk prediction, explanation and visualization in one place.
  • Modular architecture - separate, extensible modules for chronic diseases and epidemiology.
  • On-premises deployment option - data stays inside hospital or health-authority infrastructure, for full control over data.
  • Designed with a path to clinical implementation - architecture and documentation aligned with a future clinical validation pathway.

How does .be healthy work?

Step by step

  1. 1
    Input

    Data collection

    Electronic Health Records, medical registries and epidemiological sources feed the pipeline from day one.

    Output: aggregated clinical data

    This step ensures every source is captured before analysis begins.

  2. 2
    Validation

    ETL & clinical validation

    An extract-transform-load pipeline applies data quality checks and clinical validation before anything moves forward.

    Output: quality-checked, validated dataset

    This step ensures data quality issues are caught before they reach the model.

  3. 3
    Feature layer

    Feature engineering

    Raw values become clinically meaningful variables and time-series features the model can actually learn from.

    Output: clinically meaningful variables

    This step ensures raw values become clinically meaningful signals.

  4. 4
    Modeling

    AI model training & validation

    Models such as XGBoost and LightGBM are trained and evaluated against held-out data, not just training performance.

    Output: trained and evaluated models

    This step ensures the model is tested against real-world performance, not just training data.

  5. 5
    Explainability

    Explainability & clinical reports

    SHAP-based explanations and feature importance accompany every prediction the model makes.

    Output: interpretable prediction logic

    This step ensures every prediction can be traced back to the factors behind it.

  6. 6
    Output

    Dashboards, alerts & recommendations

    Risk scores, alerts and recommendations reach clinical and operational teams through interactive dashboards.

    Output: actionable decision support

    This step ensures insights reach clinical and operational teams in time to act.

Multi-language potential
English Polish Arabic German and more

AI at .be healthy

Built for clinical validation - not just accuracy

Accuracy alone is not enough in clinical decision support. Be Healthy focuses on the types of errors that matter in screening workflows - especially missed high-risk patients.

0% Sensitivity

Share of high-risk patients correctly flagged.

0% Specificity

Share of low-risk patients correctly cleared.

0% Precision

Share of flagged cases that were truly high-risk.

0 Missed high-risk cases

False negatives out of 200 illustrative cases.

Illustrative Balanced-mode snapshot from the interactive validation lab below - internal PoC / demonstration figures, not certified clinical performance metrics.

Interactive validation lab

Switch validation modes to see how the same model trades off missed high-risk cases against unnecessary alerts - the matrix and metrics below update instantly.

Screening mode Validation sample: 200 cases Clinical review
Predicted high-risk Predicted low-risk
Actual high-risk 0 True positive 0 False negative
Actual low-risk 0 False positive 0 True negative

Why false negatives matter: in clinical screening, false negatives matter because they represent patients who may need attention but were not flagged by the model.

Recall / sensitivity 0%
Specificity 0%
Precision 0%
Accuracy 0%

Clinical meaning

Prioritizes catching more high-risk patients, accepting more false alerts.

Demo values for interface explanation only. Final clinical performance requires prospective validation.

  • Explanation of each individual prediction, not just aggregate model statistics
  • Design informed by clinical guideline frameworks such as ESC/ACC/AHA
  • Compliance-aware architecture, built with GDPR/MDR-style constraints in mind
  • AI as support for doctors - not a replacement for clinical judgment

Example module

See Doctor Assistant in action

The Heart Disease Risk Assistant is one example of Doctor Assistant functionality - future clinical modules are planned for RSV, sepsis, diabetes and broader chronic disease burden.

Patient risk preview
Risk score Elevated

Main drivers

  • Age
  • Chest pain type
  • Max heart rate
  • ST depression

Recommended action

Review risk profile and consider further diagnostic pathway.

Open the full demonstrator

About this module

A high-level look at what Doctor Assistant represents as a decision-support layer, and how its explanations should be read.

What this demo illustrates

This demonstrator presents the idea of patient-level decision support with transparent risk scoring. It shows how clinical variables can be converted into an explainable risk signal for discussion, not into an autonomous diagnosis.

Explainability approach

The demo uses simplified, contribution-style explanations to show why a score changes. This is intended for interpretability in a PoC context and does not expose a production model or full clinical validation workflow.

Practical use
  • Patient-level risk support for clinician review
  • Structured, explainable discussion of contributing factors
  • A foundation for future modules across chronic disease and epidemiology
Limitations
  • Not a diagnostic tool
  • Not for independent clinical decision-making
  • Requires clinical validation
  • Uses illustrative demo logic

Use-case demonstrators

Example Modules

Illustrative, browser-only PoC demonstrators of both platform layers - no backend, no patient data sent. Each is one example of a broader module family, not the full product.

Example Module · Doctor Assistant demonstrator

Heart Disease Risk - Browser-Only PoC Demo

Adjust the patient parameters, run the demo scoring model, and inspect the SHAP-style explanation behind the resulting risk indication. No patient data is sent anywhere - everything below runs locally in your browser.

Browser-only PoC demo No backend No patient data sent SHAP-style explanation PoC benchmark only
ModeBrowser-only PoC demo
Model typeDemo scoring model (XGBoost-inspired)
ExplanationSHAP-style feature impact
BenchmarkPoC benchmark only - not certified

Patient Clinical Parameters

Demo input

Risk Assessment

-
--% Illustrative risk score (demo)
LowModerateElevated

Recommended Action

-

This is a decision-support signal, not a diagnosis. Always correlate with symptoms, history and laboratory results, and consult a qualified clinician.

SHAP-Style Feature Impact

Demo explanation
Baseline --% Final --%

Increases demo risk Protective in this demo

This explanation visualizes signed feature contributions of the browser demo scoring function. It is inspired by SHAP-style explanations, but it is not a certified clinical XAI engine.

Browser-only demonstrator. This calculator uses a transparent, rule-based scoring function - it is not the XGBoost model referenced on this site, does not diagnose disease, and must never inform real patient care. Always consult qualified healthcare professionals for medical decisions.

About this module

A high-level look at what this demonstrator represents, and the reference model it is inspired by - trained and benchmarked offline, not running in your browser.

1,025Patient records (reference dataset)
13Clinical features
0.99AUC (PoC benchmark)
What this demo illustrates

This demonstrator presents the idea of patient-level decision support with transparent risk scoring. It shows how clinical variables can be converted into an explainable risk signal for discussion, not into an autonomous diagnosis. Several tree-based models were evaluated offline on a reference dataset; the live calculator above is a transparent rule-based heuristic inspired by those results.

Data & validation context

The reference model was benchmarked on a small, single-center dataset using standard cross-validation, reaching an internal PoC AUC of around 0.99 on a held-out test set. That figure describes the offline reference model, not a certified clinical performance result - real deployment would require a larger, multi-center dataset and independent validation.

Explainability approach

The demo uses simplified, contribution-style explanations to show why a score changes - in the spirit of SHAP-style explainability, presented at a high level. This is intended for interpretability in a PoC context and does not expose a production model, its training pipeline, or a full clinical validation workflow.

Practical use
  • Patient-level risk support for clinician review
  • Transparent, explainable scoring for discussion - not an autonomous diagnosis
  • A starting point for structured conversations about risk factors
Limitations
  • Not a diagnostic tool
  • Not for independent clinical decision-making
  • Uses illustrative demo logic - reference dataset is small and single-center, with no external validation
  • Requires clinical validation before any real-world use

Example Module · Disease Forecasting demonstrator

Influenza Monitoring - Browser-Only Demonstrator

Synthetic, illustrative regional incidence data explored through an interactive command-center view. Generated in your browser - not real surveillance data, and not for public health decisions.

ModeBrowser-only PoC demo
DataSynthetic / illustrative
MethodMoving average + simple extrapolation
UseNot for public health decisions

Controls

Synthetic / illustrative data. Not for public health decisions.

Incidence Time Series

Synthetic demo data
Selected region Comparison region Forecast (demo, dashed) Peak week

Forecast (demo)

Forecast index --/100
Seasonal pressure --
Next 4 weeks (demo)
--------
Current trend -
Peak week -- -- cases / 100k
Mean vs. current -- --
Next-period direction -
Region comparison -
Weeks above threshold (120) --

Regional Status

All 16 regions, synthetic
Synthetic / illustrative data. Generated in your browser for demonstration purposes only - not real epidemiological observations, and not for public health decisions.

About this module

A high-level look at what this demonstrator represents, and the reference dataset behind the seasonal patterns it illustrates.

16Regions (reference dataset)
97Weekly observations
~2Seasons of history
What this demo illustrates

This demonstrator shows how seasonal disease-incidence signals can be translated into early planning cues for hospital and public-health management. It focuses on trend direction, seasonal timing and regional pattern comparison rather than certified epidemiological forecasting.

Data & validation context

The reference dataset behind this module contains weekly influenza incidence across 16 regions over roughly two seasons, measured as cases per 100,000 inhabitants so regions of different sizes stay comparable. It is used to illustrate seasonal patterns, such as winter peaks and inter-regional synchrony, not to represent live surveillance data. The live chart above uses separate, browser-generated synthetic data for the same reason.

Forecasting approach

Forecasting signals are generated from time-series patterns - trend, seasonality and regional correlation - validated against the historical reference data at a high level. This demo does not expose the underlying statistical modeling pipeline or parameters.

Practical use
  • Early warning before seasonal pressure increases
  • Planning beds, staff and procurement before peak demand
  • Comparing regional or scenario-level pressure patterns
  • Supporting management discussions with transparent signals
Limitations
  • PoC only - illustrative reference dataset, not live hospital or public-health data
  • Short time horizon (~2 seasons) limits long-term trend analysis
  • No weather, vaccination-rate or mobility data included
  • This page's live chart uses synthetic, randomly generated data, not the reference dataset above
  • Requires external validation before operational or public-health deployment

Knowledge & Technology

Knowledge & Technology

How the platform is architected, deployed and secured - and the statistical & machine-learning methods behind both proof-of-concept modules, in one place.

Technology

Technology - Architecture & Roadmap

How data flows through the platform, how it can be deployed under institutional control, and where the modular architecture is designed to go next.

End-to-End Pipeline

  1. Data import - EHR systems, epidemiological databases, medical registries.
  2. ETL & validation - quality checks, outlier detection, clinical validation, standardization.
  3. Feature engineering - derived features, lag variables, domain-specific transforms.
  4. Model training & validation - cross-validation, hyperparameter tuning, held-out evaluation.
  5. Explainability analysis - SHAP values, feature importance, clinical validation reports.
  6. Dashboards & alerts - interactive visualizations, thresholds and automated alerting.

System Architecture

A three-tier design intended to keep the platform modular, scalable, and easier to secure:

Presentation

  • Interactive dashboards for clinicians and administrators
  • Responsive design for desktop and mobile
  • Role-based access control

Logic

  • API for model serving and data queries
  • Model registry with versioning
  • Clinical rules and alerting thresholds

Data

  • Structured store for clinical data
  • Time-series store for epidemiological surveillance
  • Object storage for model artifacts and logs

Deployment model

On-premises deployment keeps data inside hospital or health-authority infrastructure:

  • No cloud upload of patient data by default
  • The institution retains control over security and access policy
  • Designed with GDPR-style data-protection principles in mind

Security & Compliance

Data security principles

  • Encryption at rest and in transit
  • Role-based access with audit logging
  • Anonymization / pseudonymization for analytics
  • Backup and recovery procedures

Regulatory posture

This platform concept is not a certified medical device. The architecture is designed with a future regulatory pathway in mind:

  • GDPR-aware data handling (minimization, right to erasure, consent)
  • Documentation structure aligned with a future MDR/FDA-style validation process
  • A full audit trail of predictions, access and model updates

Ethical design principles

  • Models are decision-support tools, not autonomous diagnostic systems.
  • Clinical decisions always rest with qualified healthcare professionals.
  • Bias and fairness monitoring across demographic groups is part of the intended process.

Methodology

Tabular machine learning - Heart Disease module

  • Preprocessing: quality checks, outlier analysis, feature scaling
  • Model comparison: XGBoost, LightGBM, Random Forest, Gradient Boosting, SVC
  • Stratified k-fold cross-validation
  • Evaluation focused on minimizing false negatives
  • SHAP-based explainability for clinical alignment

Time series analysis - Influenza module

  • Stationarity testing (Augmented Dickey-Fuller)
  • STL decomposition into trend / seasonal / residual components
  • ACF / PACF autocorrelation analysis and lag-feature engineering
  • K-means clustering of regional epidemic profiles
  • Classical forecasting (SARIMA, Holt-Winters) suited to limited data

Future Modules

Planned disease modules

  • Diabetes risk stratification
  • Chronic kidney disease - early detection of declining renal function
  • Sepsis early warning
  • Stroke risk prediction
  • Multi-pathogen respiratory surveillance, extending the influenza module

Advanced capabilities under consideration

  • Direct EHR integration
  • Telemedicine-facing risk scores
  • Natural language processing of clinical notes
  • Federated learning across institutions without sharing patient data

Looking for formulas and methodology deep-dives?

The Knowledge Base covers classification metrics, ROC/AUC, SHAP explainability and time-series methods behind both platform layers.

Reference

Knowledge Base

Core concepts, formulas and practical context behind the Heart Disease (classification) and Influenza (time-series) proof-of-concepts - a quick reference, not a textbook.

Overview

This knowledge base consolidates core concepts, mathematical formulas, and practical insights from the platform's two proof-of-concept modules:

  • Heart Disease module - binary classification using supervised learning
  • Influenza module - time series analysis and forecasting

Purpose

A quick reference for the statistical and machine-learning methods used across the platform. Each topic covers a clear definition, the underlying formula, and the practical context from the PoCs.

Use the panel on the right to jump between topics - each is self-contained.

Classification Metrics

Core concepts

Supervised learning: training a model on labeled data (input → output) to predict outcomes for new data.

Binary classification: predicting one of two classes (e.g. disease present = 1, absent = 0).

Confusion matrix

The foundation for all classification metrics:

  • TP (True Positive) - correctly predicted positive cases
  • TN (True Negative) - correctly predicted negative cases
  • FP (False Positive) - incorrectly predicted positive (Type I error)
  • FN (False Negative) - incorrectly predicted negative (Type II error)

Accuracy

\[ \text{Accuracy} = \frac{TP + TN}{TP + TN + FP + FN} \]

Proportion of correct predictions overall.

Precision (positive predictive value)

\[ \text{Precision} = \frac{TP}{TP + FP} \]

Of all predicted positives, how many were actually positive?

Recall (sensitivity)

\[ \text{Recall} = \frac{TP}{TP + FN} \]

Of all actual positives, how many did we correctly identify?

F1-score

\[ F1 = 2 \cdot \frac{\text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}} \]

Balances precision and recall into a single metric.

Why it matters in the Heart Disease PoC

Multiple models were compared using these metrics. High recall is critical in medical screening (missed cases are costly), but precision matters too (too many false alarms erode trust).

ROC Curve & AUC

ROC curve

A graphical plot of the trade-off between True Positive Rate (Recall) and False Positive Rate at various classification thresholds.

False Positive Rate

\[ FPR = \frac{FP}{FP + TN} \]

AUC (Area Under the Curve)

Measures the model's ability to distinguish between classes.

AUC interpretation

  • 1.0 - perfect classifier
  • 0.9-0.99 - excellent
  • 0.8-0.89 - good
  • 0.5 - random guessing

Why it matters in the Heart Disease PoC

The best-performing internal PoC model reached an AUC around 0.99 on the held-out test set - an internal benchmark result, not a certified clinical performance figure. The ROC curve helps select an operating threshold based on clinical priorities (e.g. prioritizing recall).

Explainability (Feature Importance & SHAP)

Feature importance

Tree-based models (XGBoost, Random Forest) measure how important each feature is for making predictions.

SHAP (SHapley Additive exPlanations)

SHAP values explain individual predictions by showing how each feature contributes to the output.

SHAP value definition

For a prediction \( f(x) \), the SHAP value \( \phi_i \) for feature \( i \) is:

\[ \phi_i = \sum_{S \subseteq F \setminus \{i\}} \frac{|S|! (|F| - |S| - 1)!}{|F|!} \left[ f(S \cup \{i\}) - f(S) \right] \]

Where \(F\) is the set of all features and \(S\) a subset not including feature \(i\).

SHAP visualizations

  • Summary plot - feature importance and impact across all samples
  • Force / waterfall plot - step-by-step contribution for a single prediction
  • Decision plot - how features push a prediction from baseline to output

Why it matters in the Heart Disease PoC

SHAP helps explain why the demo model flags elevated risk - e.g. "elevated ST depression, chest pain type, and exercise-induced angina drove this score." This kind of explanation is central to clinical trust and decision support.

Time Series Basics (Trend, Seasonality, Residual)

Decomposition

Additive model

\[ Y_t = T_t + S_t + R_t \]

\(T_t\) trend, \(S_t\) seasonal component, \(R_t\) residual (noise).

STL decomposition

STL (Seasonal and Trend decomposition using Loess) is a robust method that can handle changing seasonal patterns, outliers, and gaps in the data.

Why it matters in the Influenza PoC

Decomposing weekly incidence data revealed clear winter seasonal peaks; the trend component showed multi-year patterns, and residuals helped flag unusual weeks.

ACF, PACF & Lag Features

Autocorrelation function (ACF)

ACF

\[ \rho_k = \frac{\text{Cov}(Y_t, Y_{t-k})}{\text{Var}(Y_t)} \]

Partial autocorrelation (PACF)

Correlation between \(Y_t\) and \(Y_{t-k}\) after removing the effect of intermediate lags - helps identify model order.

Lag features

Lag feature

\[ \text{lag}_k(Y_t) = Y_{t-k} \]

Cross-correlation

Pearson correlation

\[ r = \frac{\sum (X_i - \bar{X})(Y_i - \bar{Y})}{\sqrt{\sum (X_i - \bar{X})^2} \sqrt{\sum (Y_i - \bar{Y})^2}} \]

Why it matters in the Influenza PoC

ACF showed strong correlation at lag = 1 week. Cross-correlation between regions such as Śląskie and Mazowieckie was high, suggesting shared nationwide dynamics. Lag features improved short-horizon forecasting.

Stationarity & ADF Test

Stationarity

A time series is stationary if its statistical properties (mean, variance, autocorrelation) remain constant over time. Many classical models (ARIMA) assume stationarity.

Augmented Dickey-Fuller (ADF) test

A statistical test for whether a series has a unit root (is non-stationary).

Interpretation

  • If p < 0.05: reject the null hypothesis → series is stationary
  • If p ≥ 0.05: fail to reject → series is non-stationary

Why it matters in the Influenza PoC

An ADF test on a reference regional series supported stationarity (p < 0.05), likely because strong seasonality removes the underlying trend - allowing classical models to be used without differencing.

Clustering (K-Means, PCA, Elbow)

K-means clustering

An unsupervised algorithm that partitions data into \(k\) clusters by minimizing within-cluster variance.

Objective (sum of squared errors)

\[ \text{SSE} = \sum_{i=1}^{k} \sum_{x \in C_i} \| x - \mu_i \|^2 \]

Elbow method

Plotting SSE vs. \(k\) - the "elbow" where SSE flattens suggests an optimal cluster count.

PCA (Principal Component Analysis)

A dimensionality-reduction technique that projects data onto orthogonal components capturing maximum variance - used here to visualize high-dimensional weekly series in 2D.

Why it matters in the Influenza PoC

K-means (k = 4) on the 16 regions revealed that provinces cluster by epidemic style (high peaks, flat patterns, frequent fluctuation) rather than by geography - PCA made this visible in two dimensions.

Practical Notes (Limitations & Next Steps)

Heart Disease PoC - takeaways

Input: 1,025 patient records, 13 clinical features. Output: binary risk prediction. Best model: XGBoost (internal PoC benchmark). Key limitations: small sample, no modern biomarkers or imaging, single-center data, no external validation.

Influenza PoC - takeaways

Input: weekly incidence, 16 regions, ~97 weeks. Methods: STL, ACF/PACF, k-means, PCA, ADF, lag features. Key limitations: short horizon, no external covariates (weather, vaccination, mobility), province-level granularity only, no severity data.

Cross-validation & overfitting

Overfitting: a model performs well on training data but poorly on new data. Cross-validation: split data into \(k\) folds, train on \(k-1\), validate on the rest, repeat.

K-fold CV score

\[ \text{CV Score} = \frac{1}{k} \sum_{i=1}^{k} \text{Score}_i \]

Next steps

  • Heart Disease: validate on external datasets, add imaging/biomarkers, explore temporal risk modeling.
  • Influenza: integrate weather and mobility data, extend to multi-pathogen surveillance.
  • General: any real deployment requires rigorous validation, regulatory compliance and integration with existing healthcare infrastructure.

Contact the .be healthy team

Discover the application of artificial intelligence in medicine

Our Team

A multidisciplinary team spanning healthcare AI, machine learning, product strategy, business development and communication - working to turn clinical data into practical, explainable decision-support tools.

  • AI & Machine Learning
  • Clinical analytics
  • Business strategy
  • Product communication

Opens your email client with a pre-filled message to contact@skillandchill.com.