You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
🩺 Machine Learning diabetes prediction model using Support Vector Machine (SVM) classifier. Analyzes 8 medical features (glucose, BMI, age, etc.) from Pima Indian dataset to predict diabetes risk with 75-80% accuracy. Built with Python, scikit-learn, pandas. Includes data preprocessing, model training, and prediction system for diabetes..
MAUDEMetrics: Open-source no-code desktop app (macOS, Windows) and Docker service for automated extraction and standardization of FDA MAUDE medical device adverse event data via the openFDA API. Bypasses the 26,000-record limit and produces reproducible, analysis-ready datasets — no programming required.
Highlighting expertise in data migration, data normalization and standardization, this project demonstrates successful data transfer from Snowflake to Databricks. It emphasizes optimized data flow and enhanced accessibility through standardization, showcasing a commitment to ethical data practices.
A Python-based data cleaning project to streamline Quickbooks invoice data for analysis, paving the way for improved insights into sales, pricing, and inventory management.
A new package processes textual descriptions of drone designs to extract structured summaries of their operational capabilities. It focuses on identifying and categorizing key features such as locomot
A machine learning model that predicts diabetes using logistic regression on medical data. The model analyzes patient features like glucose levels and BMI to classify diabetes risk.
This Data Analytics project focused on understanding the career preferences and motivations of Generation Z.Through survey data and analysis, this project aims to identify key trends and factors influencing their career choices, providing insights for employers,educators, and recruiters looking to engage with this new generation of talent.
Hi folk, During my internship at KultureHire, I completed an end to end Data Analytics project. I created an executive and functional dashboard using pivot tables, conducted a thorough analysis, and provided actionable recommendations. I'm excited to share my work and the insights I discovered.
🌟 Data Cleaning and Processing 🌟 Handled missing values, removed duplicates, standardized salary formats, and treated outliers for consistency.Revealed trends in company performance, job roles, and salary distributions after refining the dataset. This project highlights the power of data preprocessing as the backbone of reliable analytics.
Iris Species Segmentation applies and compares three unsupervised clustering algorithms — K-Means, Agglomerative Clustering, and DBSCAN — on the classic Iris dataset, using PCA for dimensionality reduction and Silhouette Score for evaluation. Includes an interactive Streamlit app for live exploration of each model's clusters and parameters.
Standardize biodiversity datasets to Darwin Core — a bilingual (PT-BR/EN) Shiny app with smart field mapping, taxonomic & coordinate validation, sensitive-species generalization, and Darwin Core Archive export.
This project is about cleaning and preparing a global layoffs dataset for analysis, focusing on handling null values, correcting data types, and ensuring data integrity for more accurate insights.
csv-managed is a Rust command-line utility for high‑performance exploration and transformation of CSV data at scale, emphasizing streaming, typed operations, and reproducible workflows via schema and index files.