NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 1 B.Tech - Computer Science / Information Science / Information Technology - Semester V NoSQL Databases A Comprehensive Undergraduate Textbook CHAPTER 1 NoSQL Fundamentals Unit I: NoSQL Fundamentals Sections Covered 1.1 Why NoSQL? 1.2 Theoretical Foundations of Distributed Data 1.3 Introduction to the Four Types of NoSQL 1.4 Strategic Database Selection 1.5 Further Reading / Viewing 1.6 Assessment Questions NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 2 Table of Contents Chapter Introduction / Overview 3 Learning Outcomes 3 Key Concepts 4 1.1 Why NoSQL? 5 The Value of Relational Databases vs. NoSQL 6 The Impedance Mismatch Problem 8 1.2 Theoretical Foundations of Distributed Data 10 The CAP Theorem 11 The BASE Properties 13 1.3 Introduction to the Four Types of NoSQL 15 Document-Oriented Databases 16 Key-Value Pairs 18 Column-Oriented Databases 19 Graph Databases 20 1.4 Strategic Database Selection 22 Factors Influencing the Choice of Database 23 Suitable Real-World Use Cases 25 1.5 Further Reading / Viewing 27 1.6 Assessment Questions 28 NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 3 Chapter 1: NoSQL Fundamentals Chapter Introduction / Overview Modern applications generate data at a speed, scale and variety that traditional single-server database designs were not originally built to handle. Social media platforms store billions of posts and interactions, e-commerce systems maintain large product catalogues with changing attributes, mobile apps create continuous event streams, and IoT devices produce high-volume sensor readings. These requirements created the need for flexible, scalable and highly available database systems called NoSQL databases. This chapter introduces the foundations of NoSQL. It begins by explaining why NoSQL emerged, while also recognising the continuing value of relational databases. It then discusses the impedance mismatch problem that occurs when object-oriented application structures are forced into relational tables. The chapter next explores the theoretical ideas behind distributed data management, especially the CAP theorem and BASE properties. Finally, it examines the four major NoSQL data models and provides a strategic framework for selecting the right database for a real-world application. By the end of this chapter, students will be able to understand not only what NoSQL is, but also when to use it, when not to use it, and how to make a justified database design decision for modern software systems. Learning Outcomes 1. Explain the need for NoSQL databases in the context of modern large-scale applications. 2. Compare the strengths and limitations of relational databases and NoSQL databases. 3. Describe the impedance mismatch problem between object-oriented programming models and relational database models. 4. Explain the CAP theorem and analyse the trade-offs among consistency, availability and partition tolerance. 5. Define the BASE properties and contrast them with the ACID properties of relational transactions. 6. Identify and differentiate the four major categories of NoSQL databases: document, key-value, column- oriented and graph databases. 7. Select an appropriate database type for a given application based on data structure, query pattern, scalability and consistency requirements. 8. Discuss suitable real-world use cases for different NoSQL database systems. Key Concepts Glossary of Key Terms in This Chapter NoSQL: A broad family of non-relational database systems designed for flexible schemas, distributed storage, horizontal scalability and high availability. Relational Database: A database system that organises data into tables with rows and columns, commonly queried using SQL and governed by relational principles. Schema: The formal structure of a database, including tables, fields, data types and relationships. In NoSQL systems, schemas are often flexible or application-defined. Horizontal Scaling: Increasing system capacity by adding more machines or nodes rather than increasing the power of a single server. Impedance Mismatch: The difficulty of mapping object-oriented application data into relational tables, especially when objects contain nested structures or relationships. Distributed Database: A database whose data is stored across multiple machines, regions or nodes while appearing as a single system to users or applications. Consistency: A property that ensures all clients see the same correct data after a successful update. Availability: A property that ensures every request receives a response, even if some nodes in the system fail. Partition Tolerance: The ability of a distributed system to continue operating despite network failures that split nodes into isolated groups. NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 4 BASE: An approach to distributed systems that favours availability and scalability through Basically Available, Soft state and Eventual consistency behaviour. Eventual Consistency: A consistency model in which replicas may temporarily differ but will converge to the same value if no new updates occur. Document Database: A NoSQL database that stores semi-structured documents, commonly in JSON or BSON format. Key-Value Store: A NoSQL database that stores data as a collection of keys mapped to values. Column-Oriented / Wide-Column Database: A NoSQL database that stores data in rows with flexible column families, suited to very large and sparse datasets. Graph Database: A NoSQL database that stores entities as nodes and relationships as edges, making it efficient for relationship-heavy queries. Polyglot Persistence: The practice of using multiple database technologies within the same application, each selected for a specific workload. 1.1 Why NoSQL? NoSQL databases emerged as a response to new data management challenges that became common with the growth of web applications, cloud computing, mobile platforms, big data and real-time analytics. The term NoSQL is commonly interpreted as "Not Only SQL", meaning that these databases do not simply reject SQL or relational databases; rather, they provide additional models and design choices for situations where relational systems may not be ideal. Figure 1.1: Evolution from relational databases to NoSQL systems For decades, relational database management systems were the dominant solution for enterprise data. They provided structured storage, powerful query capabilities, strong transaction guarantees and a mature ecosystem of tools. However, large-scale internet systems introduced requirements such as rapid schema changes, geo-distributed users, very high write volume and low-latency access. These workloads encouraged the development and adoption of NoSQL systems. Did You Know?: NoSQL does not always mean "no SQL at all". Many NoSQL databases provide SQL-like query languages, indexing and transaction features. The main difference is the underlying data model and distributed design philosophy. 1.1.1 The Value of Relational Databases vs. NoSQL Relational databases are not outdated. They remain the best choice for many applications, especially where data is highly structured and correctness is more important than flexible scaling. Banking systems, inventory management, accounting applications and enterprise resource planning systems often depend on relational databases because they require accurate transactions and well-defined relationships. NoSQL databases are valuable when application data is large, rapidly changing, semi-structured, distributed across many servers or queried in ways that do not fit naturally into tables and joins. The decision is therefore NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 5 not "SQL versus NoSQL" in a simple sense. A better question is: which database model best matches the application workload? Table 1.1: The Value of Relational Databases versus NoSQL Databases Dimension Relational Databases NoSQL Databases Data Model Tables with rows and columns; relationships expressed using keys. Documents, key-value pairs, wide columns or graphs, depending on database type. Schema Usually fixed and predefined before data insertion. Usually flexible; fields may vary between records. Query Language SQL is a standard and powerful declarative language. Varies by product; may use APIs, JSON queries, CQL, Cypher or SQL-like languages. Transactions Strong ACID transactions are a core strength. May support limited or tunable transactions; often designed around BASE principles. Scalability Traditionally vertical scaling; modern systems can also scale horizontally with complexity. Designed for horizontal scaling across many nodes. Joins Strong support for joins across related tables. Joins are often avoided; data is frequently denormalised for query speed. Consistency Strong consistency is common by default. Consistency may be strong, eventual or tunable depending on system design. Best Use Cases Financial transactions, ERP, payroll, structured reporting. Large catalogues, real-time sessions, social feeds, IoT logs, recommendations, distributed apps. Examples MySQL, PostgreSQL, Oracle Database, SQL Server. MongoDB, Redis, Cassandra, HBase, Neo4j, DynamoDB. Relational databases are especially strong when the data model is stable, the application needs complex joins, and transactions must be correct under all circumstances. NoSQL databases are especially strong when the data model changes frequently, when data must be distributed across many servers, or when very high read/write throughput is required. Table 1.2: Choosing SQL or NoSQL by Application Requirement Requirement Relational Database Advantage NoSQL Advantage Bank transfer Debiting one account and crediting another must be atomic and strongly consistent. Not usually preferred unless the NoSQL system supports strong multi-document transactions. Product catalogue Works well if all products have similar attributes. Works better when every product category has different attributes such as size, colour, processor, lens or warranty. User session cache Possible, but may be slower and unnecessarily structured. Key-value stores provide extremely fast lookup and expiry of session data. Social network relationships Many joins may become complex and expensive. Graph databases naturally store friends, followers, paths and communities. IoT sensor stream May be difficult at huge write volumes. Wide-column databases handle large, distributed and time-series workloads efficiently. Key Insight: A relational database is usually a safe default for structured transactional data. A NoSQL database becomes attractive when the application demands flexible structure, large-scale distribution, high availability or relationship-specific traversal that is difficult to express with relational joins. 1.1.2 The Impedance Mismatch Problem The impedance mismatch problem refers to the conceptual and technical difficulty of converting data between two different representations: object-oriented programming structures used in applications and table-based structures used in relational databases. In programming languages such as Java, Python or C#, developers often work with objects that contain nested objects, lists and complex relationships. In relational databases, data must be decomposed into tables, rows, columns and foreign keys. For example, an e-commerce application may represent an order as a single object containing customer details, shipping address, payment status and a list of purchased items. A relational database usually stores this NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 6 information across several tables such as Orders, Customers, Addresses, OrderItems and Payments. The application must then use joins or object-relational mapping tools to reconstruct the original object. Table 1.3: Examples of Object-Relational Impedance Mismatch Mismatch Area Object-Oriented Application View Relational Database View Typical Difficulty Identity Objects have identity in memory. Rows are identified by primary keys. Object identity and database identity must be mapped correctly. Structure Objects may contain nested objects and lists. Tables are flat structures with rows and columns. Nested data must be split across multiple tables. Relationships References can directly connect objects. Foreign keys and join tables represent relationships. Complex relationships require multiple joins. Inheritance Classes may inherit from other classes. Tables do not naturally support inheritance. Developers must choose mapping strategies. Schema Evolution Object fields can change during development. Database schema changes require migrations. Frequent changes can slow agile development. Query Style Application navigates objects. Database uses set-based SQL queries. Developers must translate between two thinking styles. Document databases reduce this mismatch for many applications because they allow related information to be stored together as a single document. Instead of spreading an order across many tables, the application can store one order document containing nested item details. This design can simplify development and improve read performance when the application usually retrieves the whole order together. { "order_id": "ORD1025", "customer": { "name": "Ananya Rao", "email": "ananya@example.com" }, "shipping_address": { "city": "Bengaluru", "pincode": "560001" }, "items": [ {"product": "Laptop", "qty": 1, "price": 62000}, {"product": "Mouse", "qty": 2, "price": 800} ], "payment_status": "Paid" } The above JSON-like document resembles the way an application naturally represents an order object. However, this does not mean document databases are always better. If an application frequently needs complex joins, strict constraints and multi-record transactions, a relational database may still be the better choice. 1.2 Theoretical Foundations of Distributed Data Many NoSQL databases are distributed systems. This means their data is spread across multiple machines to improve scalability, performance and fault tolerance. Distributed data management is powerful, but it introduces challenges that do not occur in a single-machine database. Machines can fail, networks can become slow, and different copies of the same data may temporarily disagree. The theoretical foundations of distributed data help database designers understand these trade-offs. Two of the most important concepts for NoSQL systems are the CAP theorem and the BASE properties. CAP explains what cannot be achieved perfectly at the same time during network failures. BASE explains a design approach that favours availability and eventual convergence in large distributed systems. NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 7 1.2.1 The CAP Theorem (Consistency, Availability, Partition Tolerance) The CAP theorem states that in a distributed data system, when a network partition occurs, the system cannot simultaneously guarantee both strong consistency and full availability. Because partitions are unavoidable in real distributed networks, a distributed database must make a trade-off: it can preserve consistency by rejecting or delaying some requests, or it can preserve availability by responding to requests even if some replicas may be temporarily out of date. Figure 1.2: CAP theorem trade-off among Consistency, Availability and Partition Tolerance The three properties are defined as follows: Consistency: Every read receives the most recent successful write or an error. All users observe the same correct value. Availability: Every request receives a non-error response, even if some replicas or nodes are unavailable. Partition Tolerance: The system continues to operate despite communication failures between groups of nodes. In a partition-free situation, a distributed system may appear to provide both consistency and availability. The important lesson of CAP is what happens during a partition. Since modern applications often run across networks, cloud regions and clusters, partition tolerance is usually considered mandatory. The practical choice becomes CP or AP behaviour. Table 1.4: Practical Interpretation of CAP Trade-offs CAP Choice Meaning During a Partition Typical Behaviour Suitable Examples CP Consistency and Partition Tolerance are prioritised. Some requests may fail or wait so that stale data is not returned. Banking ledger, inventory reservation, strongly consistent configuration store. AP Availability and Partition Tolerance are prioritised. System continues responding, but some users may temporarily see old data. Social media feeds, likes, shopping recommendations, user activity streams. CA Consistency and Availability are prioritised only when partitions are not considered. Not realistic for a truly distributed system during network partition. Single-node relational database or tightly controlled local cluster. Tunable Application chooses consistency level per operation. Reads and writes can be configured for stronger or weaker consistency. Cassandra-style systems where quorum settings influence behaviour. NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 8 Common Misconception: CAP does not say that a database can have only two properties forever. It says that during a network partition, a distributed system must choose between strong consistency and full availability. Consider a distributed shopping application deployed in two regions. If the network link between the regions fails, customers in both regions may continue placing orders. An AP design accepts orders in both regions and resolves conflicts later. A CP design may stop accepting some orders until it can verify the latest stock count. The first approach improves availability, while the second protects correctness. 1.2.2 The BASE Properties BASE is a design philosophy commonly associated with highly available distributed NoSQL systems. It stands for Basically Available, Soft state and Eventual consistency. Unlike ACID, which emphasises immediate correctness of transactions, BASE accepts temporary inconsistency in order to provide high availability and scalability. Figure 1.3: ACID and BASE transaction mindsets Basically Available: The system attempts to respond to every request, even if some data is temporarily stale or some functionality is degraded. Soft State: The state of the system may change over time even without direct user input because background synchronisation and replication are occurring. Eventual Consistency: If no new updates are made, all replicas will eventually converge to the same value. Table 1.5: Comparison of ACID and BASE Approaches Property ACID-Oriented Systems BASE-Oriented Systems Main Goal Correct and reliable transactions. High availability and scalability. Consistency Strong and immediate consistency. Often eventual or tunable consistency. State Database state changes only through committed transactions. State may change as replicas synchronise in the background. Failure Handling May reject or roll back transactions to preserve correctness. May continue serving requests with temporary inconsistency. Best For Payments, accounting, inventory locks, legal records. Feeds, likes, logs, recommendations, sessions, distributed content. Typical Databases PostgreSQL, Oracle, MySQL, SQL Server. Cassandra, DynamoDB, Couchbase, Riak-style systems. Eventual consistency is acceptable when temporary differences are not harmful. For example, if one user sees 10 likes on a post and another user sees 11 likes for a short time, the application can still function. However, NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 9 eventual consistency may be unacceptable for a bank balance or exam result update, where users expect immediate correctness. Exam Tip: Use ACID when correctness must be immediate. Use BASE when high availability, scale and tolerance of temporary inconsistency are more important. 1.3 Introduction to the Four Types of NoSQL NoSQL is not a single database model. It is a category of database systems that includes multiple approaches to storing and querying data. The four major types are document-oriented databases, key-value stores, column- oriented or wide-column databases, and graph databases. Each type is designed for a different kind of data structure and access pattern. Figure 1.4: Four major types of NoSQL databases Understanding these four types is essential because a wrong NoSQL choice can create as many problems as a wrong relational design. For example, using a key-value store for complex relationship queries may lead to inefficient application code, while using a graph database for simple caching may be unnecessarily complex. 1.3.1 Document-Oriented Databases A document-oriented database stores data as documents, commonly represented in JSON, BSON or XML-like formats. A document is a self-contained record that may include nested fields, arrays and varying structures. Documents are grouped into collections, which are similar to tables but usually do not require every document to have the same fields. Document databases are especially useful when application objects naturally contain nested information. Product catalogues, user profiles, content management systems, medical records and order management systems often benefit from document storage because related information can be kept together. { "student_id": "CS501", "name": "Ravi Kumar", "semester": 5, "skills": ["SQL", "Python", "MongoDB"], "address": {"city": "Mysuru", "state": "Karnataka"}, "projects": [ {"title": "Library App", "technology": "Flask"}, {"title": "Student Analytics", "technology": "Pandas"} ] } NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 10 Strengths: Flexible schema, natural JSON-style representation, easy storage of nested data, rapid application development and good read performance for document-level queries. Limitations: Complex joins are not as natural as relational databases; duplicated data can create update challenges; document size and design must be planned carefully. Common Examples: MongoDB, CouchDB, Couchbase, Amazon DocumentDB. 1.3.2 Key-Value Pairs A key-value database is the simplest form of NoSQL database. Each item is stored as a unique key associated with a value. The database does not need to understand the internal structure of the value; it only retrieves, stores or deletes the value using its key. This simplicity allows extremely fast access. A key-value store is similar to a dictionary or hash map in programming. For example, a login session can be stored using the session ID as the key and the session details as the value. When the user sends a request, the application can quickly retrieve the session using the key. Table 1.6: Simple Examples of Key-Value Storage Key Value Use session:9a71 {user_id: 105, role: student, expires: 10:30} User login session cart:U102 [PEN10, BOOK22, BAG05] Temporary shopping cart otp:779921 473819 One-time password cache page:home <html>cached homepage</html> Page cache Strengths: Very fast read and write operations, simple design, excellent for caching, sessions, counters and temporary data. Limitations: Poor support for complex queries; values are often opaque to the database; relationships and filtering must be handled by application logic. Common Examples: Redis, Amazon DynamoDB, Riak, Memcached. 1.3.3 Column-Oriented Databases Column-oriented NoSQL databases, often called wide-column databases, store data in rows and column families. Unlike a relational table, each row does not need to have the same columns. This makes the model suitable for very large, sparse and distributed datasets. Wide-column databases are designed for high write throughput and horizontal scalability. A wide-column database should not be confused with analytical columnar warehouses. In NoSQL, the term usually refers to systems such as Cassandra and HBase, where data is organised by row keys and column families. The design is heavily influenced by the expected query pattern. Table 1.7: Conceptual View of Wide-Column Storage Row Key Column Family: Profile Column Family: Activity Possible Use user_101 name=Kiran, city=Hubballi login_2026_06_20=true, clicks=85 User behaviour tracking sensor_220 location=PlantA, type=temperature 2026-06-20T10:00=31.2, 10:01=31.4 IoT time-series storage device_14 model=X2, owner=Lab1 error_501=voltage_low, error_502=fan_stop Machine monitoring Strengths: Massive write scalability, efficient storage of sparse data, high availability across clusters and strong support for time-series or event-style workloads. Limitations: Data modelling is query-driven and can be difficult for beginners; ad-hoc joins are not natural; consistency may be tunable rather than always strong. Common Examples: Apache Cassandra, Apache HBase, Google Bigtable, ScyllaDB. NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 11 1.3.4 Graph Databases A graph database stores data as nodes, edges and properties. Nodes represent entities such as people, products, accounts or cities. Edges represent relationships such as FRIEND_OF, PURCHASED, TRANSFERRED_TO or LOCATED_IN. Properties store additional details about nodes or edges. Graph databases are powerful when the relationships between data items are as important as the data items themselves. Instead of performing many joins, a graph database directly traverses relationships. This makes it suitable for social networks, recommendation engines, fraud detection, network analysis, knowledge graphs and route planning. Table 1.8: Basic Elements of a Graph Database Graph Element Meaning Example Node An entity in the system. Student, Course, Teacher, Account, Product Edge A relationship between two nodes. ENROLLED_IN, TEACHES, FRIEND_OF, BOUGHT Property Additional information attached to a node or edge. Student name, course credits, purchase date, transfer amount Traversal Moving through connected nodes and edges. Find friends of friends; detect circular money transfer paths Strengths: Excellent relationship traversal, intuitive modelling of connected data, efficient path and network queries, useful for recommendations and fraud patterns. Limitations: Not ideal for simple bulk key lookup or very large append-only event logs; graph modelling requires careful understanding of relationships. Common Examples: Neo4j, JanusGraph, Amazon Neptune, TigerGraph, ArangoDB. Table 1.9: Comparison of the Four Types of NoSQL Databases Type Data Model Best For Example Query / Need Common Databases Document JSON/BSON documents with nested fields. Flexible records, catalogues, profiles, content. Find products where brand is X and category is mobile. MongoDB, CouchDB, Couchbase Key-Value Unique key mapped to a value. Caching, sessions, counters, shopping carts. Get session data for session ID S123. Redis, DynamoDB, Riak Wide-Column Rows with column families and flexible columns. Massive write workloads, IoT, time-series, logs. Fetch sensor readings for device D between two times. Cassandra, HBase, Bigtable Graph Nodes, edges and properties. Relationship-heavy and path- based queries. Find shortest connection between two users. Neo4j, JanusGraph, Neptune 1.4 Strategic Database Selection A database should never be chosen only because it is popular or new. Strategic database selection means choosing a database based on the application requirements, data structure, query patterns, consistency needs, expected scale, development team skills and operational constraints. A good database choice makes the application simpler, faster and more reliable. A poor choice may create performance issues, complex code and long-term maintenance problems. NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 12 Figure 1.5: Strategic database selection framework 1.4.1 Factors Influencing the Choice of Database The following factors should be analysed before selecting a database for a project. Data Structure: Is the data tabular, nested, key-based, sparse or highly connected? Query Pattern: Will the application use joins, document retrieval, key lookup, time-range scans or relationship traversal? Transaction Requirement: Does the system require strict ACID transactions or can it tolerate eventual consistency? Scalability: Is the application expected to run on one server, a cluster or multiple regions? Read/Write Ratio: Is the workload read-heavy, write-heavy or balanced? Latency Requirement: Does the system need millisecond response, real-time streaming or batch analytics? Availability Requirement: Must the system remain available even during node failures or network partitions? Schema Flexibility: Will the data fields change frequently during product development? Security and Compliance: Does the application handle regulated data such as payment, health or identity information? Team Skill and Ecosystem: Does the team have experience with the database, monitoring tools, backup strategy and deployment model? Cost and Operations: What are the licensing, cloud, hardware, maintenance and scaling costs? Integration Needs: Does the database integrate well with existing applications, analytics pipelines and reporting tools? Table 1.10: Mapping Requirements to Database Choices Requirement / Situation Recommended Database Direction Reason Strict financial transaction Relational database or strongly consistent distributed SQL. ACID correctness and constraints are critical. Rapidly changing product attributes Document database. Flexible schema can store different structures per product category. Login sessions and cache Key-value store. Fast lookup using a unique key and support for expiry. Billions of sensor readings Wide-column database. Designed for high write volume and distributed time-series access. Fraud ring detection Graph database. Relationships and paths are central to the problem. Business dashboards over structured data Relational database / data warehouse. SQL and structured reporting are strong requirements. NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 13 Global social feed AP-style NoSQL database. Availability and large-scale distribution are more important than immediate consistency. Mixed application with payments, catalogue and recommendations Polyglot persistence. Different components require different data models and guarantees. Design Principle: Start from the queries, not from the database brand. A database is suitable only when its data model supports the most important queries efficiently and safely. 1.4.2 Suitable Real-World Use Cases The suitability of a database becomes clearer when viewed through realistic application scenarios. The same organisation may use more than one database because each workload has a different requirement. Table 1.11: Suitable Real-World Use Cases for Database Selection Domain / Application Suitable Database Type Why It Fits Illustrative Example E-commerce Product Catalogue Document Database Products have different attributes across categories; documents store nested information naturally. A laptop has RAM and processor; a shoe has size and material; both can be documents in one collection. Shopping Cart and Session Store Key-Value Store Fast retrieval by cart ID or session ID; data is temporary and simple. cart:user105 -> list of products selected by the user. IoT Sensor Monitoring Wide-Column Database Large-scale, write-heavy time-series data with range queries by device and time. Store temperature readings from thousands of sensors every second. Social Network Graph Database Relationships such as follows, friends, likes and communities are core to the application. Find mutual friends or recommend connections. Banking Core Ledger Relational Database Strong transactions, constraints and auditability are mandatory. Transfer money from account A to account B atomically. Recommendation Engine Graph or Document Database Recommendations depend on relationships or flexible user/product profiles. Users who bought X also bought Y; people similar to you watched Z. Log Analytics Platform Wide-Column or Search-Oriented NoSQL High write volume, distributed storage and time-based access. Search error logs from a production application. Content Management System Document Database Articles, pages and metadata vary in structure and can be stored as documents. Store a blog post with author, tags, images and comments. Leaderboards and Counters Key-Value Store Counters can be incremented quickly with low latency. Game score ranking or likes count. Fraud Detection in Finance Graph Database Suspicious patterns often appear as hidden relationship chains. Detect accounts connected through shared devices, phone numbers or transfer paths. Many modern systems use polyglot persistence. For example, an e-commerce platform may use a relational database for payments, a document database for product catalogues, a key-value store for user sessions, a graph database for recommendations and a wide-column database for clickstream analytics. This approach increases architectural complexity but allows each component to use the most suitable storage model. Table 1.12: Polyglot Persistence Example for an E-Commerce Platform System Component Possible Database Choice Reason Payment and invoices Relational database Requires ACID transactions and audit records. Product catalogue Document database Product attributes vary by category and change frequently. User sessions Key-value store Needs low-latency lookup and automatic expiry. User behaviour events Wide-column database Handles large write-heavy event streams. Recommendations Graph database Uses relationships among users, products and purchases. Business reporting Data warehouse / relational analytics store Supports SQL reporting and BI dashboards. NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 14 Practical Checklist: Before selecting a NoSQL database, write the top five queries the application must answer, estimate the expected data volume and growth, identify the required consistency level, and decide whether the team can operate the system safely in production. 1.5 Further Reading / Viewing The following resources are recommended for students who wish to study NoSQL and distributed data systems in greater depth. Foundational Textbooks Sadalage, P. J., & Fowler, M. (2012). NoSQL Distilled: A Brief Guide to the Emerging World of Polyglot Persistence. Addison-Wesley. - A clear beginner-friendly explanation of NoSQL concepts and database selection. Kleppmann, M. (2017). Designing Data-Intensive Applications. O'Reilly Media. - A highly respected book on distributed data, replication, partitioning and consistency. Tanenbaum, A. S., & Van Steen, M. Distributed Systems. - Useful for understanding distributed systems, failures and communication models. Date, C. J. An Introduction to Database Systems. - A classic reference for relational database principles. Important Papers and Concepts Codd, E. F. (1970). A Relational Model of Data for Large Shared Data Banks. - Introduced the relational model. Brewer, E. A. (2000). Towards Robust Distributed Systems. - Popularised the CAP theorem idea. Gilbert, S., & Lynch, N. (2002). Brewer's Conjecture and the Feasibility of Consistent, Available, Partition- Tolerant Web Services. - Formalised CAP theorem arguments. Chang, F. et al. (2008). Bigtable: A Distributed Storage System for Structured Data. - Influential work behind wide-column distributed storage. Official Documentation and Platforms MongoDB Documentation - document database concepts, schema design and aggregation framework. Redis Documentation - key-value storage, caching, data structures and pub/sub use cases. Apache Cassandra Documentation - wide-column modelling, partitioning and tunable consistency. Neo4j Documentation - graph data modelling and Cypher query language. AWS DynamoDB Documentation - managed key-value and document database design patterns. Suggested Videos / Online Learning Introductory videos on CAP theorem and distributed systems. MongoDB University courses for document database fundamentals. DataStax Academy resources for Cassandra and distributed NoSQL concepts. Neo4j GraphAcademy courses for graph database modelling. Cloud provider tutorials on selecting databases for application workloads. 1.6 Assessment Questions Part A - Multiple Choice Questions (1 Mark Each) Q1. The term NoSQL is commonly interpreted as: 1. No Structured Query Language 2. Not Only SQL 3. No Software Query Layer NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 15 4. New Object SQL Answer: (b) Q2. Which database model is usually best for strict ACID transactions such as bank transfers? 1. Key-value store 2. Graph database 3. Relational database 4. Document database Answer: (c) Q3. The impedance mismatch problem mainly occurs between: 1. Programming objects and relational tables 2. CPU and RAM 3. Indexes and queries 4. Cloud and local storage Answer: (a) Q4. In the CAP theorem, P stands for: 1. Performance 2. Partition Tolerance 3. Persistence 4. Parallelism Answer: (b) Q5. Which CAP behaviour prioritises consistency during a network partition? 1. AP 2. CP 3. CA 4. BASE Answer: (b) Q6. BASE stands for: 1. Basic Access, Safe Execution 2. Basically Available, Soft state, Eventual consistency 3. Binary Allocation Storage Engine 4. Balanced Availability and Strong Execution Answer: (b) Q7. Which NoSQL type stores data as JSON-like records? 1. Document database 2. Graph database 3. Column database 4. Key-only database Answer: (a) Q8. Which database type is most suitable for friend-of-friend queries? NoSQL Fundamentals - Chapter 1 Chapter 1: NoSQL Fundamentals | Unit I: NoSQL Fundamentals | Page 16 1. Key-value store 2. Graph database 3. Relational spreadsheet 4. File system Answer: (b) Part B - Short Answer Questions (5 Marks Each) Q9. Explain why NoSQL databases became important for modern web-scale applications. Q10. Compare relational databases and NoSQL databases using any five dimensions. Q11. What is the impedance mismatch problem? Explain with a simple e-commerce order example. Q12. Define the three components of the CAP theorem and explain why all three cannot be fully guaranteed during a network partition. Q13. Explain BASE properties and compare them with ACID properties. Q14. Write short notes on document databases and key-value stores with suitable examples. Q15. Differentiate between column-oriented NoSQL databases and graph databases. Part C - Long Answer / Essay Questions (10 Marks Each) Q16. Discuss the need for NoSQL databases. In your answer, explain the limitations of relational databases in large-scale distributed applications and the advantages offered by NoSQL systems. Q17. Explain the CAP theorem in detail. Use examples to show the difference between CP and AP systems. Why is partition tolerance important in distributed databases? Q18. Describe the four major types of NoSQL databases. For each type, explain its data model, strengths, limitations and suitable real-world use cases. Q19. Explain strategic database selection. What factors should be considered before choosing a database for a software project? Support your answer with examples. Part D - Analytical / Case-Based Questions Q20 (Case Study). A start-u