← Back to projects

PUBLISHED ON GITHUB

Heart Disease Data Analysis

An exploratory analysis of demographic and clinical variables associated with heart disease presence. The published repository documents data cleaning, visualizations and analytical limitations.

PythonPandasNumPyMatplotlibJupyter
UCI Heart Disease · Cleveland subset303 SOURCE RECORDS
14VARIABLES
01

Dataset

303 records and 14 original variables from the Cleveland subset of UCI Heart Disease.

02

Data cleaning

Missing-value markers were normalized. The ca and thal variables were converted to numeric types and missing values imputed using the mode. Duplicate checks found no duplicates.

03

Python analysis

Pandas and Matplotlib were used to examine sex, age, chest pain type and maximum heart rate. No predictive machine learning model was developed.

04

Reported insights

The repository reports heart disease presence in approximately 55.3% of male patients and 25.8% of female patients in this sample. These are descriptive associations within the dataset.

Have a question worth exploring?

Let’s work together