INTRODUCTION TO DATA SCIENCE • B.E/ B.Tech 5th Semester • CSE / IS / IT Unit 1 Foundations of Data Science Sections: Evolution · Lifecycle · DS vs AI · Big Data · BI · Analytics · Roles · Applications · Challenges Sections 1.1 – 1.6 Level Semester V Type Theory + Case Studies Prepared by Ashwini S | Introduction t o Data Science Unit 1 — Sections at a Glance Six core sections that build from historical context to practical roles and future opportunities 1.1 Evolution & Significance Historical milestones from 1663 to 2023+; why DS is central to every industry 1.2 Data Science Lifecycle 8-phase iterative process: problem framing → collection → cleaning → EDA → modelling → evaluation → deployment → monitoring 1.3 DS vs Artificial Intelligence Nested relationship between AI, ML, DL, and DS; key differences in goals, methods, tools, and outputs 1.4 Big Data, BI & Analytics 5 V's of Big Data; technologies (Hadoop, Spark, Kafka); BI vs DS; four types of analytics 1.5 Components & Data Scientist Role Five core pillars of DS; end-to-end responsibilities from problem framing to ethical deployment 1.6 Analyst, Engineer — Apps & Challenges Role comparisons; applications across 9 domains; technical, organisational & ethical challenges; emerging opportunities What is Data Science? A foundational definition before we explore its evolution, lifecycle, and applications Definition: Data Science is an interdisciplinary field that combines statistics, mathematics, computer science, and domain expertise to extract meaningful knowledge, patterns, and actionable insights from structured and unstructured data using scientific methods, algorithms, and systems. What makes it 'interdisciplinary'? • Statistics & Probability – foundation of inference and uncertainty • Computer Science – algorithms, programming, data structures • Mathematics – linear algebra, calculus, optimisation • Domain Knowledge – healthcare, finance, engineering context • Communication – turning insights into business decisions • Ethics & Law – responsible data handling (GDPR, DPDP Act) What does a Data Scientist do? • Collects, cleans, and prepares raw data for analysis • Explores data visually and statistically (EDA) • Builds and trains predictive models (ML/DL) • Deploys models into real-world production systems • Monitors model performance over time (MLOps) • Communicates findings to technical & non- technical teams Why is Data Science important now? • 2.5 quintillion bytes of data created every day (IBM) • Affordable cloud computing makes large-scale analysis feasible • Advances in ML algorithms (especially deep learning) • Every industry from healthcare to agriculture now generates data • Evidence-based decision-making replaces intuition-based choices • WEF: Data roles among top 10 fastest-growing jobs globally Evolution of Data Science — Part 1 (1663–1989) From early statistics and probability to the first formal data mining frameworks 1663 John Graunt – Bills of Mortality Analysed London death records to find patterns in plague mortality. First known quantitative population data analysis — laid the groundwork for modern statistics, epidemiology, and public health analytics. 1805–1830 Legendre & Gauss — Least Squares Independently developed the method of least squares for fitting data to models. This remains the foundational regression technique in modern predictive modelling and machine learning today. 1936 Alan Turing — Turing Machine Defined the theoretical model of computation. Set the mathematical foundation for all algorithmic data processing, programming languages, and modern computers that power Data Science. 1960s Database Systems Emerge Development of hierarchical (IBM IMS) and network database models enabled systematic data storage and retrieval. Shift from paper records to digital structured data — the precursor to modern SQL databases. 1974 Peter Naur Coins 'Data Science' In 'Concise Survey of Computer Methods', Naur proposed Data Science as an alternative to computer science — focused on data processing as a scientific discipline. First formal academic use of the term. 1977 John Tukey — Exploratory Data Analysis Published 'Exploratory Data Analysis', shifting focus from confirmatory hypothesis testing to open-ended, visual, data-driven discovery. EDA remains a core phase of every modern Data Science lifecycle. 1989 Knowledge Discovery in Databases (KDD) The first KDD workshop was held at IJCAI, formally introducing the process of extracting knowledge from large databases. Set the stage for the data mining era and the later CRISP-DM lifecycle model. Evolution of Data Science — Part 2 (1996–2023+) From data mining to Big Data, deep learning, AutoML, and the era of Generative AI 1996 First KDD Conference Establishment of data mining as a discipline with emphasis on automated pattern recognition in large datasets. Foundations of supervised and unsupervised learning as practical tools, not just theory. 2001 W.S. Cleveland – 'An Action Plan' Proposed expanding statistics into a broader field called Data Science, incorporating computing. Identified six components: multidisciplinary investigations, models & methods, computing, pedagogy, tool evaluation, theory. 2006 Apache Hadoop Released Enabled distributed storage (HDFS) and processing (MapReduce) of massive datasets across commodity hardware. Democratised Big Data analytics — organisations no longer needed supercomputers for large-scale analysis. 2008 'Data Scientist' Title Coined DJ Patil (LinkedIn) and Jeff Hammerbacher (Facebook) coined the professional title 'Data Scientist'. Triggered industry-wide adoption. Patil later became the first US Chief Data Scientist (2015–2017). 2012 Harvard Business Review Article 'Data Scientist: The Sexiest Job of the 21st Century' by Davenport & Patil. Mainstream recognition of DS as a high- demand profession, triggering global talent searches and university curriculum changes. 2014–2016 Rise of Deep Learning & Cloud GPU-powered neural networks achieved human-level performance in image recognition, speech, and NLP. AWS, Azure, GCP democratised massive-scale compute — ML became accessible to any organisation. 2017–2022 AutoML, MLOps, Explainable AI AutoML tools automated model selection and hyperparameter tuning. MLOps formalised CI/CD pipelines for ML. XAI frameworks (LIME, SHAP) addressed the 'black box' problem in regulated industries. 2023+ Large Language Models & Generative AI GPT-4, Gemini, Llama and other LLMs integrate into DS workflows — automated code generation, data analysis, natural language queries. Generative AI reshapes how data scientists work and communicate insights. Significance of Data Science Why DS has become indispensable across economy, science, government, and careers Economic Significance • Organisations using DS consistently outperform competitors in revenue growth and cost reduction • Amazon: recommendation engine drives ~35% of total revenue • Netflix: 80% of content watched is driven by the recommendation algorithm • Uber: real-time surge pricing and driver dispatch uses live ML models • DS enables demand forecasting, supply chain optimisation, and dynamic pricing • McKinsey: data-driven companies are 23× more likely to acquire customers Scientific & Research Significance • Genomics: DS enables sequencing and analysing millions of genomes in days • Physics: CERN generates ~15 PB of data/year from particle collision experiments • Climate: DS models are used to simulate global temperature changes • Drug discovery: AlphaFold (DeepMind) predicted 3D structure of 200M+ proteins • Epidemiology: COVID-19 contact tracing, vaccine distribution models, ICU forecasting • Astronomy: ML identifies exoplanets in Kepler telescope data at scale Social & Governance Significance • Smart city planning: traffic, energy, waste management optimised with real-time data • Public health: disease surveillance systems for early outbreak detection • Tax & finance: fraud detection systems in revenue departments • Education policy: dropout prediction, student performance analytics • Census & resource allocation: evidence-based policy formulation • Disaster response: predictive flood and wildfire risk mapping Career & Industry Demand • World Economic Forum: Data Analyst and Data Scientist in top 5 emerging roles (2023) • LinkedIn: 'Data Scientist' one of the top 3 fastest-growing job titles globally • US BLS: 35% projected growth in DS roles between 2022–2032 (much faster than average) • Average salary: USD 120,000+ in the US; ₹12–25 LPA entry-level in India • Demand exceeds supply by 50% — talent shortage persists across geographies • DS skills are valued in every sector: IT, banking, pharma, logistics, government Data Science Lifecycle — Overview CRISP-DM inspired 8-phase iterative process: insights at any phase may require revisiting earlier steps 1 Business Understanding Define problem & KPIs 2 Data Collection Databases, APIs, IoT, scraping 3 Data Cleaning Handle noise, missing, duplicates 4 EDA Visualise, summarise, find patterns 5 Feature Eng. & Modelling Train & tune ML models 6 Model Evaluation Metrics, test sets, overfitting check 7 Deployment APIs, pipelines, dashboards 8 Monitoring Drift detection, retraining The lifecycle is ITERATIVE — poor model performance (Phase 6) may require revisiting Feature Engineering (Phase 5) or even Data Collection (Phase 2). This is normal and expected. ⟳ Data Science Lifecycle — Phases 1 to 4 From understanding the business problem through to exploratory data analysis Phase 1: Business / Problem Understanding • Most critical phase — a poorly framed problem wastes entire project effort • Collaborate with domain experts, managers, and end-users to clarify objectives • Translate vague business goals into precise, measurable data science questions • Define success metrics: e.g., 'reduce churn by 15%' or 'achieve >90% recall' • Conduct feasibility analysis — is the required data available and accessible? • Deliverable: Formal problem statement, project charter, stakeholder agreement • Common pitfall: Starting with data instead of the problem leads to irrelevant models Phase 2: Data Collection & Acquisition • Identify all potential data sources: internal databases, third-party APIs, surveys, logs • Structured sources: SQL/relational databases, Excel, CSV, ERP systems • Unstructured sources: web scraping, PDFs, social media, images, IoT sensor streams • Document data lineage and provenance — where does each field come from? • Understand data access constraints: privacy rules, licensing, sampling biases • Deliverable: Raw data repository with documented sources, formats, and volumes • Key question: Is the data sufficient in quantity, quality, and time-coverage? Phase 3: Data Cleaning & Preprocessing • Real-world data is always dirty — studies show 60–80% of project time spent here • Missing values: deletion (if <5%), mean/median/mode imputation, predictive imputation • Outlier detection: Z-score, IQR method, DBSCAN; decide to remove, cap, or keep • Duplicates: exact and near-duplicate records must be identified and merged or dropped • Encoding: one-hot encoding for nominal, label encoding for ordinal variables • Normalisation (Min-Max) and Standardisation (Z-score) for numerical stability • Deliverable: Clean, consistently formatted dataset ready for EDA and modelling Phase 4: Exploratory Data Analysis (EDA) • Goal: understand data distributions, detect anomalies, discover relationships • Descriptive statistics: mean, median, mode, standard deviation, skewness, kurtosis • Univariate analysis: histograms, box plots, density plots for each feature • Bivariate analysis: scatter plots, correlation matrices (Pearson, Spearman) • Identify class imbalance in the target variable — critical for classification tasks • Tools: Python (Pandas, Matplotlib, Seaborn), R (ggplot2), Tableau, Jupyter • Deliverable: EDA report with distribution plots, correlation heatmap, key insights Data Science Lifecycle — Phases 5 to 8 From model building and evaluation through to production deployment and continuous monitoring Phase 5: Feature Engineering & Modelling • Feature engineering: creating new features from raw data that better capture patterns • Techniques: interaction terms, polynomial features, binning, TF-IDF, word embeddings • Dimensionality reduction: PCA (for numerical), LDA (for classification tasks) • Algorithm selection: depends on problem type — regression, classification, clustering • Bias-variance trade-off: underfitting (high bias) vs overfitting (high variance) • Cross-validation: k-fold CV ensures robust performance estimates on limited data • Deliverable: Trained model pipeline with documented architecture and hyperparameters Phase 6: Model Evaluation & Validation • Test only on held-out test set — never use test data during training or tuning • Regression metrics: MAE (mean absolute error), RMSE, MAPE, R² (explained variance) • Classification metrics: Accuracy, Precision, Recall (Sensitivity), F1-Score, AUC-ROC • Confusion matrix: TP, TN, FP, FN — essential for understanding error types • Clustering metrics: Silhouette Score, Davies-Bouldin Index, Calinski-Harabasz • Hyperparameter tuning: Grid Search, Random Search, Bayesian optimisation (Optuna) • Deliverable: Performance report; final model chosen based on business success metrics Phase 7: Deployment & Communication • A model that isn't deployed has zero business value — deployment bridges DS and engineering • Batch deployment: scheduled model predictions (e.g., nightly churn scores) • Real-time deployment: REST APIs (Flask, FastAPI) respond to live prediction requests • Containerisation: Docker packages model + dependencies for portable, consistent execution • Orchestration: Kubernetes manages scaling of deployed models under load • Model registry & versioning: MLflow tracks experiments and production model versions • Communication: Tableau/Streamlit dashboards, executive reports, explaining model logic Phase 8: Monitoring & Maintenance • Data drift: statistical distribution of input features changes over time (e.g., user behaviour) • Concept drift: relationship between features and target changes (e.g., COVID disrupting demand) • Monitoring tools: Evidently AI, Arize AI, Fiddler — alert on performance degradation • Retraining strategy: trigger-based (when accuracy drops below threshold) or scheduled • MLOps: CI/CD pipelines for ML automate testing, validation, and deployment of new models • A/B testing: compare production model vs new challenger model on live traffic • Deliverable: Monitoring dashboard, SLA alerts, retraining schedule, version audit trail Model Evaluation — Key Metrics by Problem Type Choosing the right metric is as important as choosing the right algorithm Problem Type Metric Formula / Definition When to Use Regression MAE Mean of |yᵢ − ŷᵢ| When all errors are equally important Regression RMSE √( Mean of (yᵢ−ŷᵢ)² ) When large errors should be penalised heavily Regression R² 1 − SS_res / SS_tot To explain variance explained by the model Classification Accuracy Correct / Total predictions When classes are balanced Classification Precision TP / (TP + FP) When false positives are costly (spam detection) Classification Recall TP / (TP + FN) When false negatives are costly (cancer diagnosis) Classification F1-Score 2×(P×R)/(P+R) When you need balance between Precision and Recall Classification AUC-ROC Area under ROC curve; 1.0 = perfect For comparing classifiers regardless of threshold Clustering Silhouette Score Range: −1 to +1; higher = better clusters Measuring cluster cohesion and separation Data Science vs Artificial Intelligence — Conceptual Relationship Often used interchangeably, but they are distinct disciplines with different goals and scope Artificial Intelligence (AI) Broadest concept: building systems that simulate human intelligence — reasoning, perception, language, action Machine Learning (ML) Subset of AI: systems learn from data using statistical algorithms without explicit programming Deep Learning (DL) Subset of ML: multi-layer neural networks model highly complex patterns in images, text, audio Data Science Intersects AI/ML/DL but also includes data engineering, domain knowledge, statistical analysis, and communication All ML-based AI uses Data Science pipelines — both fields are needed When a data scientist builds a churn model using a neural network, they practise both Data Science (lifecycle pipeline) and AI (learning system). Not all Data Science uses AI A data analyst creating a monthly revenue dashboard in SQL and Power BI is practising Data Science with zero AI — purely descriptive analytics. Not all AI uses Data Science A symbolic AI chess engine (like Deep Blue, 1997) uses hand-crafted rules and search trees — no training data, no data pipeline, no Data Science lifecycle. Data Science vs AI — Detailed Comparison Eight key dimensions distinguishing how the two fields differ in goal, scope, and practice Dimension Data Science Artificial Intelligence Primary Goal Extract insights, patterns & knowledge from data to support human decision-making Build systems that autonomously perform tasks requiring human- like intelligence Core Focus Data collection, cleaning, EDA, modelling, and communication of findings Reasoning, learning, language understanding, perception, planning, and autonomous action Output Dashboards, reports, predictive models, segmentation, recommendations Autonomous systems, intelligent agents, generative models, robotic controllers Data Dependency Always depends on historical data — no data means no Data Science May use data (ML-based) or explicit rules (symbolic AI like Prolog or expert systems) Human Involvement High — humans interpret results, validate models, and make final decisions Varies from high (supervised ML) to fully autonomous (self-driving vehicles) Scope Includes non-ML work: SQL dashboards, BI, statistical reports, surveys Broader in goal (autonomy & intelligence); may not involve data pipelines at all Key Tools Python, R, SQL, Pandas, Scikit-learn, Tableau, Power BI, Jupyter TensorFlow, PyTorch, OpenAI API, Hugging Face, Prolog, ROS (Robot OS) Real Examples Customer churn prediction, fraud score, demand forecast, A/B test GPT-4, AlphaGo, Tesla Autopilot, DeepMind AlphaFold, Amazon Alexa Big Data & the 5 V's Framework Datasets so large and complex that traditional tools fail — defined by five key characteristics IBM estimates 2.5 quintillion bytes (2.5 × 10¹⁸) of data are created every single day. Big Data refers to datasets whose Volume, Velocity, and Variety exceed conventional processing capabilities. Managing and extracting value from such data requires specialised frameworks and technologies. Volume • Scale: data ranging from terabytes (TB) to petabytes (PB) and exabytes (EB) • Facebook: processes 100+ TB of interaction data per day • NASA generates ~2.4 TB/day from Earth observation satellites • Challenge: Traditional RDBMS cannot store or query data at this scale • Solution: Distributed file systems (HDFS), object storage (AWS S3, GCS) Velocity • Speed at which data is generated, captured, processed, and acted upon • NYSE generates approximately 1 TB of new trade data every trading day • Twitter: over 500,000 tweets are posted every minute worldwide • Spectrum: batch (hourly/nightly) → micro-batch → real-time streaming • Solution: Apache Kafka (streaming), Apache Flink, Spark Streaming Variety • Structured: relational tables with defined schema (SQL, CSV, Excel) • Semi-structured: JSON, XML, YAML — partial schema, nested fields • Unstructured: text, social media posts, images, audio, video (~80% of data) • Hospital data example: EHR (structured) + X-rays (images) + doctor notes (text) • Challenge: Schema mapping, multi-modal pipelines, storage incompatibility Veracity • Trustworthiness, accuracy, and reliability of data — signal vs noise • Social media data: bots, spam, duplicate accounts distort real-world signals • IoT sensors: calibration errors, packet loss cause gaps and noise • Survey data: social desirability bias, non-response bias, recall errors • Solutions: Data quality frameworks, duplicate detection, anomaly flagging Value • Most important V — large data is worthless without extracting actionable insights • Netflix recommendation engine: drives ~80% of all content watched on platform • Value requires asking the right questions AND having a pipeline to answer them • Business value: cost savings, revenue growth, risk mitigation, process automation • Key lesson: Data is a liability (cost, risk) until transformed into Value Big Data Technologies A specialised ecosystem of frameworks for storing, processing, and streaming large-scale data Apache Hadoop • Distributed Storage (HDFS): splits files into 128 MB blocks, replicates 3× across nodes • Distributed Processing (MapReduce): splits computation into Map (filter/transform) and Reduce (aggregate) tasks • YARN: Yet Another Resource Negotiator — manages cluster resources across applications • Best suited for batch processing of large, immutable datasets • Written in Java; originally developed at Yahoo in 2005, donated to Apache • Limitation: disk-based I/O makes it slow for iterative algorithms (e.g., ML training) Apache Spark • In-memory distributed processing — stores intermediate results in RAM, not disk • Up to 100× faster than Hadoop MapReduce for iterative machine learning algorithms • APIs in Python (PySpark), Scala, Java, R — widely used for DS at scale • Spark MLlib: built-in ML library for classification, clustering, recommendation • Spark Streaming: real-time stream processing using micro-batch architecture • GraphX: graph computation engine for social network and recommendation analysis Apache Kafka • Distributed event streaming platform — acts as a high- throughput message broker • Producers publish messages (events) to topics; Consumers subscribe and process them • Retains messages for a configurable period — consumers can replay history • Handles millions of events per second with sub-10ms latency at scale • Use cases: real-time fraud detection, log aggregation, IoT telemetry pipelines • Kafka Connect: pre-built connectors to databases, cloud storage, and APIs NoSQL Databases • Designed for flexible schema, horizontal scaling, and high-throughput reads/writes • MongoDB (Document): stores JSON-like BSON documents; flexible schema; rich query API • Cassandra (Wide-Column): optimised for time-series, IoT, and write-heavy workloads • Neo4j (Graph): models relationships as edges; ideal for fraud networks, social graphs • Redis (Key-Value): in-memory store; used for caching, sessions, real-time leaderboards • When to use NoSQL: when data is unstructured, schema evolves, or scale is critical Cloud Data Warehouses • Google BigQuery: serverless, petabyte-scale SQL analytics; pay-per-query pricing • Amazon Redshift: columnar storage, massively parallel processing (MPP) architecture • Snowflake: separates storage and compute; multi- cloud; supports structured + semi-structured • All support standard SQL, making them accessible to analysts and data engineers • Data Lakehouse: hybrid architecture combining data lake flexibility + warehouse performance • Tools: dbt (data transformation), Airflow (pipeline orchestration), Fivetran (connectors) Business Intelligence vs Data Science Both use data to support decisions — but differ fundamentally in time horizon, methods, and output Dimension Business Intelligence (BI) Data Science Primary Question What happened? What is happening right now? Why did it happen? What will happen? What should we do? Time Orientation Retrospective — historical and current state analysis Forward-looking — predictive and prescriptive future analysis Analytics Type Descriptive analytics (reports) + basic Diagnostic analytics Predictive + Prescriptive analytics using ML and optimisation Primary Users Business managers, C-suite, operations teams, analysts Data scientists, ML engineers, product managers, researchers Data Types Primarily structured, clean, relational (SQL) data Structured, semi-structured, and unstructured data (text, images) Methods OLAP cubes, dashboards, KPIs, drill-down reports, data warehousing ML/DL models, EDA, statistical testing, A/B experiments Key Tools Power BI, Tableau, MicroStrategy, SAP BusinessObjects, Looker Python, R, TensorFlow, Scikit-learn, Apache Spark, Jupyter Output Scorecards, dashboards, visualisations, ad-hoc reports Predictive models, automated decisions, NLP pipelines, insights Skill Emphasis SQL, data warehousing, ETL, report design, domain knowledge Statistics, ML theory, programming, feature engineering, MLOps Time to Value Days to weeks — relatively fast to build and update Weeks to months — iterative experimentation and validation required Four Types of Data Analytics A spectrum from simple description to intelligent prescription — increasing in complexity and value 1. Descriptive 2. Diagnostic 3. Predictive 4. Prescriptive Descriptive Analytics • Core Question: What happened in the past? • Summarises historical data to describe what has already occurred • Techniques: Aggregation, pivot tables, time-series charts, histograms, frequency tables • Tools: SQL GROUP BY, Excel PivotTables, Tableau dashboards, Power BI reports • Example 1: Monthly sales report showing total revenue per product category per region • Example 2: Website analytics dashboard showing page views, bounce rate, session duration • Limitation: Only tells you WHAT happened, not WHY or what to do next Diagnostic Analytics • Core Question: Why did it happen? • Drills into data to find root causes and contributing factors • Techniques: Root cause analysis, drill-down, correlation analysis, hypothesis testing • Statistical tests: chi-square, t-test, ANOVA for comparing group differences • Example 1: Investigating WHY Q3 sales dropped 15% — broken down by region, product, rep • Example 2: Analysing which customer segments have highest churn and why • Limitation: Still backward-looking — explains past events but doesn't predict future Predictive Analytics • Core Question: What is likely to happen in the future? • Uses ML models trained on historical data to forecast future outcomes • Algorithms: Linear/logistic regression, Random Forest, XGBoost, LSTM for time series • Output is probabilistic — never certain; always carries uncertainty and confidence intervals • Example 1: Predicting which customers will churn in the next 30 days • Example 2: Forecasting product demand to optimise inventory levels for Black Friday • Requires: Clean historical data, defined target variable, cross- validation for reliability Prescriptive Analytics • Core Question: What action should we take to achieve the best outcome? • Most advanced type — recommends optimal decisions, not just predictions • Techniques: Linear/integer programming, simulation (Monte Carlo), reinforcement learning • Combines predictive models with optimisation to evaluate multiple possible actions • Example 1: Google Maps recommending the optimal route considering real-time traffic • Example 2: Determining optimal drug dosage protocol based on patient-specific predictors • Emerging: LLM-powered agents that take autonomous actions based on business objectives Core Components of Data Science Five intersecting pillars — the ideal Data Scientist spans all; in practice, teams specialise Statistics & Mathematics • Probability & distributions: Normal, Binomial, Poisson — models real-world uncertainty • Inferential statistics: hypothesis testing, p-values, confidence intervals • Linear algebra: matrices, eigenvectors, SVD — underpins PCA and neural networks • Calculus: derivatives and gradient descent drive all neural network optimisation • Bayesian reasoning: update beliefs with evidence — fundamental to probabilistic ML • Optimisation theory: convexity, Lagrange multipliers, stochastic gradient descent Programming & Computer Science • Python: primary language — NumPy, Pandas, Scikit- learn, Matplotlib, TensorFlow • R: preferred for statistical computing, academia, bioinformatics • SQL: essential for querying relational databases — JOINs, window functions, CTEs • Data structures: arrays, hash maps, graphs — critical for algorithm efficiency • Software engineering: version control (Git), modular code, unit testing • CS fundamentals: time/space complexity, recursion, APIs, RESTful services Machine Learning & AI • Supervised learning: classification (SVM, RF, XGBoost) and regression (Linear, Ridge) • Unsupervised learning: K-Means, DBSCAN, PCA, autoencoders • Deep learning: CNNs (images), RNNs/LSTMs/Transformers (sequences/text) • Natural Language Processing: tokenisation, word embeddings, BERT, GPT fine-tuning • Computer Vision: object detection (YOLO), segmentation, image classification • Reinforcement Learning: agent-environment interaction — AlphaGo, robotics, control Data Engineering & Infrastructure • ETL/ELT pipelines: extract from source, transform to clean format, load to warehouse • SQL databases: PostgreSQL, MySQL, Snowflake for structured analytics • NoSQL: MongoDB, Cassandra, Redis for flexible schema and high-throughput writes • Big Data: Hadoop (batch), Spark (in-memory), Kafka (streaming), Airflow (orchestration) • Cloud platforms: AWS (S3, Redshift, SageMaker), GCP (BigQuery, Vertex AI), Azure ML • Data lakehouse: Delta Lake, Apache Iceberg — unified batch + streaming architecture Domain Knowledge & Communication • Industry understanding: frame meaningful, commercially relevant data questions • Regulatory awareness: GDPR (EU), CCPA (California), India DPDP Act 2023 • Ethical competency: fairness, bias detection, model transparency (XAI) • Data storytelling: translate complex model outputs into clear business narratives • Visualisation tools: Matplotlib, Seaborn, Plotly, Tableau, Power BI, Streamlit • Stakeholder communication: executive presentations, technical documentation, dashboards