Predicting Healthcare Costs Using Demographic and Behavioral Factors

Author

Oluwatobi A. Alawode

execute: echo: true message: false warning: false cache: true

Project Overview

This project demonstrates how healthcare costs can be predicted using demographic and behavioral data in a way that directly mirrors real-world healthcare analytics workflows. The analysis prioritizes interpretability, transparency, and business relevance, which are critical for payer, provider, consulting, and population health analytics roles.

The objective is to identify key drivers of medical spending, evaluate model performance, and translate statistical findings into actionable insights relevant to healthcare decision-making.

Data Description

The dataset contains individual-level health insurance data with the following variables:

  • Age: Age in years
  • Sex: Male or female
  • BMI: Body Mass Index
  • Children: Number of dependents
  • Smoker: Smoking status
  • Region: Geographic region
  • Charges: Annual healthcare costs (outcome variable being predicted)

Data Preparation

Code
# Load required packages
library(tidyverse)
library(caret)
library(car)
library(broom)
library(modelsummary)
Code
# Load data
data <- read.csv("C:/Users/alawo/Downloads/archive(5)/insurance.csv")

# Remove duplicates
data <- distinct(data)

# Convert categorical variables to factors
data <- data %>%
  mutate(across(c(sex, smoker, region), factor))

Exploratory Data Analysis

Distribution of Healthcare Costs

Healthcare costs are right-skewed, indicating that a small proportion of individuals account for very high expenditures.

Smoking Status and Costs

Smokers incur substantially higher healthcare costs than non-smokers across the distribution.

Age and Healthcare Costs

Healthcare costs increase steadily with age, reflecting greater healthcare utilization over the life course.

Modeling Approach

The data were split into training (80%) and testing (20%) sets to evaluate out-of-sample performance.

Code
set.seed(123)
train_index <- createDataPartition(data$charges, p = 0.8, list = FALSE)
train_data <- data[train_index, ]
test_data <- data[-train_index, ]

A multiple linear regression model was estimated using demographic, behavioral, and regional predictors.

Code
model <- lm(charges ~ age + sex + bmi + children + smoker + region,
            data = train_data)

Model Results

Code
modelsummary(model, statistic = "std.error")
(1)
(Intercept) -10863.218
(1090.604)
age 241.783
(13.213)
sexmale -151.560
(368.333)
bmi 323.355
(31.695)
children 530.935
(153.980)
smokeryes 24396.386
(455.432)
regionnorthwest -732.727
(527.208)
regionsoutheast -876.206
(520.105)
regionsouthwest -1234.634
(530.972)
Num.Obs. 1072
R2 0.761
R2 Adj. 0.759
AIC 21703.3
BIC 21753.0
Log.Lik. -10841.631
F 422.728
RMSE 5970.13

Plain Language Interpretation

  • Smoking status is the strongest predictor of healthcare costs.
  • Older age is associated with higher medical spending.
  • Higher BMI increases healthcare expenditures.
  • Sex, region, and number of children have smaller effects.

Model Diagnostics

             GVIF Df GVIF^(1/(2*Df))
age      1.019976  1        1.009938
sex      1.011536  1        1.005752
bmi      1.104556  1        1.050979
children 1.005021  1        1.002507
smoker   1.012385  1        1.006173
region   1.093195  3        1.014962

Variance Inflation Factors indicate no serious multicollinearity.

Model Performance

Code
# Generate predictions
test_data$predicted <- predict(model, test_data)

# Performance metrics
rmse <- RMSE(test_data$predicted, test_data$charges)
mae <- MAE(test_data$predicted, test_data$charges)
r2 <- cor(test_data$predicted, test_data$charges)^2

tibble(RMSE = rmse, MAE = mae, R_squared = r2)
# A tibble: 1 × 3
   RMSE   MAE R_squared
  <dbl> <dbl>     <dbl>
1 6397. 4291.     0.704

The model performs well for low to moderate costs, with reduced accuracy for extreme high-cost cases.

Business Impact and Healthcare Analytics Relevance

This analysis highlights how behavioral risk factors drive healthcare spending. Smoking is the dominant cost driver, suggesting that targeted prevention and cessation programs could yield substantial cost savings. Age and BMI further emphasize the importance of early intervention and population health management.

From a healthcare analytics perspective, this project demonstrates end-to-end cost modeling, including data preparation, exploratory analysis, predictive modeling, validation, and interpretation for decision-makers.

Conclusion

This project demonstrates how interpretable statistical models can be used to identify healthcare cost drivers and support prevention-focused strategies. The workflow and insights align closely with real-world healthcare analytics roles in payers, providers, consulting firms, and public health organizations.