Code
# Load required packages
library(tidyverse)
library(caret)
library(car)
library(broom)
library(modelsummary)execute: echo: true message: false warning: false cache: true
This project demonstrates how healthcare costs can be predicted using demographic and behavioral data in a way that directly mirrors real-world healthcare analytics workflows. The analysis prioritizes interpretability, transparency, and business relevance, which are critical for payer, provider, consulting, and population health analytics roles.
The objective is to identify key drivers of medical spending, evaluate model performance, and translate statistical findings into actionable insights relevant to healthcare decision-making.
The dataset contains individual-level health insurance data with the following variables:
# Load required packages
library(tidyverse)
library(caret)
library(car)
library(broom)
library(modelsummary)# Load data
data <- read.csv("C:/Users/alawo/Downloads/archive(5)/insurance.csv")
# Remove duplicates
data <- distinct(data)
# Convert categorical variables to factors
data <- data %>%
mutate(across(c(sex, smoker, region), factor))Healthcare costs are right-skewed, indicating that a small proportion of individuals account for very high expenditures.
Smokers incur substantially higher healthcare costs than non-smokers across the distribution.
Healthcare costs increase steadily with age, reflecting greater healthcare utilization over the life course.
The data were split into training (80%) and testing (20%) sets to evaluate out-of-sample performance.
set.seed(123)
train_index <- createDataPartition(data$charges, p = 0.8, list = FALSE)
train_data <- data[train_index, ]
test_data <- data[-train_index, ]A multiple linear regression model was estimated using demographic, behavioral, and regional predictors.
model <- lm(charges ~ age + sex + bmi + children + smoker + region,
data = train_data)modelsummary(model, statistic = "std.error")| (1) | |
|---|---|
| (Intercept) | -10863.218 |
| (1090.604) | |
| age | 241.783 |
| (13.213) | |
| sexmale | -151.560 |
| (368.333) | |
| bmi | 323.355 |
| (31.695) | |
| children | 530.935 |
| (153.980) | |
| smokeryes | 24396.386 |
| (455.432) | |
| regionnorthwest | -732.727 |
| (527.208) | |
| regionsoutheast | -876.206 |
| (520.105) | |
| regionsouthwest | -1234.634 |
| (530.972) | |
| Num.Obs. | 1072 |
| R2 | 0.761 |
| R2 Adj. | 0.759 |
| AIC | 21703.3 |
| BIC | 21753.0 |
| Log.Lik. | -10841.631 |
| F | 422.728 |
| RMSE | 5970.13 |
GVIF Df GVIF^(1/(2*Df))
age 1.019976 1 1.009938
sex 1.011536 1 1.005752
bmi 1.104556 1 1.050979
children 1.005021 1 1.002507
smoker 1.012385 1 1.006173
region 1.093195 3 1.014962
Variance Inflation Factors indicate no serious multicollinearity.
# Generate predictions
test_data$predicted <- predict(model, test_data)
# Performance metrics
rmse <- RMSE(test_data$predicted, test_data$charges)
mae <- MAE(test_data$predicted, test_data$charges)
r2 <- cor(test_data$predicted, test_data$charges)^2
tibble(RMSE = rmse, MAE = mae, R_squared = r2)# A tibble: 1 × 3
RMSE MAE R_squared
<dbl> <dbl> <dbl>
1 6397. 4291. 0.704
The model performs well for low to moderate costs, with reduced accuracy for extreme high-cost cases.
This analysis highlights how behavioral risk factors drive healthcare spending. Smoking is the dominant cost driver, suggesting that targeted prevention and cessation programs could yield substantial cost savings. Age and BMI further emphasize the importance of early intervention and population health management.
From a healthcare analytics perspective, this project demonstrates end-to-end cost modeling, including data preparation, exploratory analysis, predictive modeling, validation, and interpretation for decision-makers.
This project demonstrates how interpretable statistical models can be used to identify healthcare cost drivers and support prevention-focused strategies. The workflow and insights align closely with real-world healthcare analytics roles in payers, providers, consulting firms, and public health organizations.