
A robust data engineering tool developed to handle missing data through custom K-means clustering and automated classification, optimizing predictive accuracy across varied datasets using Python.
This project addressed a critical challenge in data science: maintaining dataset integrity when faced with incomplete information. I engineered a two-stage pipeline—Missing Value Estimation and Optimal Classification—to automate the cleaning and predictive analysis of high-dimensional datasets. The software architecture was designed to perform sophisticated data preprocessing and model selection using a combination of Pandas, NumPy, and Scikit-learn.
Data imputation was approached through two distinct methodologies to compare the impact of local vs. global mean substitution.
df.fillna(df.mean()) to replace null values with the column-wide average.pickKCenters) and iteratively update (pickNewCenters) centroids based on Euclidean distance.Before classification, raw data undergoes a rigorous cleaning and transformation phase:
NaN to prevent skewing the statistical mean.df.dropna() for validation sets and ensured consistent data types across the dataframe to support mathematical operations.The final stage of the pipeline automates the selection of the most accurate machine learning model for a given dataset.
RandomForestClassifier).cross_validation.train_test_split, the tool measures the accuracy scores of both models. The testClassifiers function programmatically selects the "Optimal Classifier" (optClassifier) based on which method achieved the highest precision for that specific data distribution.writePredictions utility to streamline the deployment of the selected model, converting results to optimized lists for file output.np.sqrt(np.power(...).sum())), the project demonstrates a deep understanding of the vector math required for custom machine learning implementations beyond standard library calls.