DATA PRIVACY AND DATA SECURITY Unit II: Disaster Recovery and Fault Tolerance Strategic Planning, Resilience Mechanisms, and Cryptographic Detection Techniques UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Learning Outcomes and Unit Roadmap • Planning for the Worst • Creating a Backup Strategy • Designing for Fault Tolerance • Antivirus Software • Antispyware • Typical signature • Byte Streams Checksums • Custom Check sums • Cryptographic Hashes • Advanced Signatures • Fuzzy Hashing • Graph-Based Hashes for Executable Files UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Why Plan for the Worst? Defining a Disaster Any event that causes a disruption beyond normal operating procedures , requiring extraordinary measures to restore stability and service. Natural Earthquakes, floods, storms, wildfires. Cyber/Human Ransomware, sabotage, human error. Technical Power grid failure, hardware crashes. Health Pandemics, mass workforce absence. Supply Chain Third-party vendor or logistics failure. Critical Planning Principles Correlated Failures: Planning must account for events that trigger secondary failures (e.g., a flood causing a power outage). Failure of the Recovery Mechanism: What happens if the backup systems or the DRP itself fails during the incident? The "Blast Radius": Understanding the geographic and operational span of a single event. THE RECOVERY PARADOX A recovery plan is only a "plan" until it is tested. An untested recovery mechanism is often the first thing to fail when a real disaster strikes. Key Insight: Resilience requires planning for the failure of your backup plan. UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Business Continuity vs. Disaster Recovery STRATEGIC ALIGNMENT ERM Enterprise Risk Management BCP Business Continuity Planning DRP Disaster Recovery Planning DRP is a technical subset of the broader BCP framework. Feature Business Continuity (BCP) Disaster Recovery (DRP) Scope Organization-wide; business processes and people. IT systems, data, networks, and infrastructure. Goal Continuous operations and organizational resilience. Restoration of technical assets after a disaster. Timing Proactive and ongoing during disruption. Reactive; initiated post- disruption. Ownership Executive Leadership & Business Unit Heads. IT Management, CIO, & Technical Teams. Output Crisis management & relocation plans. Technical runbooks & data restore procedures. Core Distinction: BCP asks "How do we keep the doors open?", while DRP asks "How do we get the servers back online?" UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Business Impact Analysis: The Planning Foundation The Six-Step BIA Process 01 Scope Definition Establish boundaries and prioritize specific organizational units. 02 Information Gathering Collect data via surveys, interviews, and automated discovery. 03 Critical-Function Identification Distinguish essential operations from supportive tasks. 04 Impact Quantification Assign monetary and non-monetary values to potential losses. 05 Recovery-Objective Definition Set specific time and data loss targets (RTO/RPO). 06 BIA Report Formalize findings to guide executive recovery strategy decisions. Impact Dimensions Financial Lost revenue, idle labor costs, and contractual penalties. Customer Service disruption, loss of trust, and churn. Regulatory Legal fines, compliance breaches, and litigation risk. Operational Supply chain breaks and inability to deliver core services. Reputational Long-term brand damage and market value decline. Key Teaching Message: Without a rigorous BIA foundation, all subsequent recovery decisions (RTO/RPO) become mere guesswork and lack strategic alignment. UNIT II: DISASTER RECOVERY & FAULT TOLERANCE RTO, RPO, MTD, and WRT RTO & RPO CORE METRICS RTO Recovery Time Objective: Maximum acceptable duration of an outage. RPO Recovery Point Objective: Maximum acceptable data loss (measured in time). MTD Max Tolerable Downtime: Absolute limit before irreversible harm occurs. WRT Work Recovery Time: Time needed to validate data and resume normal ops. RTO + WRT ≤ MTD The Fundamental Recovery Constraint RTO asks: "How quickly must we recover?" RPO asks: "How much data can we afford to lose?" Last Backup DISASTER Systems Up Full Ops MTD (Maximum Tolerable Downtime) RPO Window RTO (Restoration) WRT (Validation) UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Availability and Recovery Tiers SYSTEM AVAILABILITY FORMULA Availability = MTBF (MTBF + MTTR) Criticality Drives Architecture: As RTO and RPO objectives tighten, the complexity and cost of the recovery solution increase exponentially. Tier classification ensures resources are allocated to the most vital business functions. Recovery Tier Classification RTO (Target) RPO (Target) Tier 1 Mission-Critical Functions < 1 Hour Near-Zero / Minutes Tier 2 Business-Critical Functions 4 – 8 Hours 1 – 4 Hours Tier 3 Operational / Essential 24 – 48 Hours Up to 24 Hours Tier 4 Non-Critical / Support 72+ Hours Varies (Days) * MTBF: Mean Time Between Failures | MTTR: Mean Time To Repair. Note: Tier 0 typically represents continuous availability (Fault Tolerance). UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Recovery Strategies and DRP Lifecycle Initiate BIA Select Strategy Develop Plan Implement Test Maintain Site Type Characteristics & Readiness Relative Cost Hot Site Fully configured; mirrored data; ready in minutes/hours. Highest Warm Site Equipped with hardware; requires data restoration. Medium Cold Site Shell space; power/cooling only; long setup time. Lowest Cloud / DRaaS Virtualized; flexible; pay-per-use; high scalability. Variable Reciprocal / Mobile Mutual aid agreements or trailer-based mobile units. Low-Medium DRP Testing Ladder Checklist: High-level review of plan components. Tabletop: Role-play scenario in a meeting room. Parallel: Systems run at DR site without interruption. Partial Interruption: Test critical systems only. Full Interruption: Entire site cutover (highest risk). "An untested DRP is a liability." Confidence in recovery requires validation under pressure. UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Historical Case: Maersk NotPetya Recovery (2017) INFECTION SCOPE & IMPACT 45,000 PCs Offline 4,000 Servers Lost 130 Countries Affected 10 Days To Rebuild USD 300 Million Revenue Impact RECOVERY TIMELINE June 2017: The Outage NotPetya ransomware propagates globally, encrypting entire Active Directory (AD) infrastructure in minutes. The Discovery IT finds a single surviving domain controller in Ghana, saved by a timely local power outage that kept it offline. The Restoration Hard drive flown from Lagos to London to serve as the master seed for global network reconstruction. The "Ghana Miracle" "The only reason we were able to recover was because of a power cut in Ghana. That domain controller was the only surviving copy of our Active Directory data." STRATEGIC SECURITY LESSONS Offline AD Backups Keep immutable, air-gapped copies of critical identity infrastructure. Network Segmentation Prevent lateral movement of malware across global geographic sites. Geographic Isolation Validate that "global" systems have localized resilience points. Supply Chain Vigilance NotPetya entered via a compromised software update (M.E.Doc). UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Backup Strategy Begins with RPO and RTO Definition: A backup is a separate, independent copy of data that enables restoration following loss, corruption, or destruction. It is the final safety net for business continuity. The Strategic Balance Frequency: How often snapshots are taken (Hourly vs. Daily). Retention: How long historical versions are kept (Compliance). Media & Location: Cloud, Tape, or Disk; On-site vs. Off-site. Security: Encryption at rest and in transit; Access control. Primary Drivers RECOVERY POINT OBJECTIVE (RPO) Determines Frequency . "How much data can we afford to lose?" (e.g., 1 hour of transactions). RECOVERY TIME OBJECTIVE (RTO) Determines Restore Speed . "How quickly must we be back online?" (e.g., 4 hours). CRITICAL DISTINCTION: A strategy must ensure not just that a "Backup Exists" but that "Recovery Works" within RTO/RPO limits. UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Backup Types Compared Backup Type Data Scope Backup Window Restore Path Storage Impact Full All selected data and files in the source. Longest Shortest Single step restoration Highest Incremental Changes since any last backup (full or incremental). Shortest Longest Full + every incremental Lowest Differential All changes since the last Full backup. Grows daily. Moderate Fast Full + latest differential Moderate Synthetic Server-side merge of full plus existing incrementals. Fast* *Server-side processing Shortest Restores as a full file High Mirror Real-time exact copy. Ransomware propagates Immediate Instant No versioning/history Highest UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Backup Schedules and Restore Paths Full + Daily Incremental Sunday: Full Backup of all selected data. Daily: Captures changes since any previous backup. Fastest backup window; minimal storage growth. Requires complex "chain" of all files for restoration. Restore Chain (Wednesday Recovery): SUN FULL MON TUE WED Full + Daily Differential Sunday: Full Backup of all selected data. Daily: Captures changes since last Full backup. Daily backup size grows as the week progresses. Only two components needed for any restoration. Restore Chain (Wednesday Recovery): SUN FULL WED DIFF Metric Incremental Strategy Differential Strategy Backup Speed Very Fast (smallest daily delta) Moderate (grows daily) Restore Speed Slowest (must process every incremental) Fast (Full + 1 Differential) Storage Usage Lowest (no redundancy between backups) Higher (redundant data in daily diffs) Continuous Data Protection (CDP) / Near-CDP Supports minute or second-level RPO by capturing every data change/snapshot. Ideal for mission-critical databases where data loss must be near-zero. UNIT II: DISASTER RECOVERY & FAULT TOLERANCE The 3-2-1 and 3-2-1-1-0 Rules The Classic 3-2-1 Rule 3 Three Copies of Data Keep at least one primary copy and two additional backups to mitigate the risk of a single point of failure. 2 Two Different Media Types Store backups on different technologies (e.g., Disk + Tape, or NAS + Cloud) to avoid correlated hardware failures. 1 One Copy Off-site Geographic separation protects against site-wide disasters like fire, flood, or regional power outages. The Modern 3-2-1-1-0 Rule 1 One Offline or Immutable Copy Maintain an Air-Gapped or WORM (Write Once, Read Many) copy that ransomware cannot encrypt or delete. 0 Zero Unverified Errors Automated recovery verification. A backup is not a backup until it has been successfully restored and validated. Adapts to modern threats where attackers actively target online backup repositories to prevent recovery. Strategic Resilience: Combating Ransomware Modern backup architecture must prioritize immutability . Air-gapped backups (physically disconnected) or logical WORM storage ensure that even if administrative credentials are compromised, the historical data remains untouchable. Verification (the "0" in 3-2-1-1-0) ensures the restore path is functional within RTO constraints. UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Backup Media and Location Decisions Storage Media Comparison Media Type Key Characteristics Restore Speed Magnetic Tape Low cost/GB, high capacity, offline storage (air-gap). Best for retention. Very Slow HDD / SSD Direct access, NAS/SAN integration. Vulnerable to ransomware if online. Very Fast Cloud Object Infinite scalability, managed by 3rd party. Off-site by design. Moderate WORM / Optical Write-Once-Read-Many. Immutable data protection. High durability. Slow Location & Disaster Risk Location Primary Benefits Site Disaster On-Site LAN speed restoration, immediate physical access. No bandwidth cost. Zero Protection Vaulted Off-site Physical isolation from regional disasters. High security. High Protection Cloud / DRaaS Geographic diversity, accessible from any internet connection. High Protection Hybrid Local cache for speed + Cloud sync for geographic safety. Balanced SECURITY DECISION Select Media based on RTO/RPO and Cost ; Select Location based on Site Disaster Risk and Ransomware Reachability . Always maintain one offline (air-gapped) copy to ensure recovery from destructive cyber-attacks. UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Backup Verification: Trust but Verify 5-Item Critical Verification Checklist Continuous Job Monitoring Review logs daily for partial failures, skipped files, or timeout errors. File-Level Restoration Sampling Perform random restore tests of individual files to ensure media readability. Full-System Restoration Exercise Validate the entire restore chain (Full + Differentials/Incrementals). Cryptographic Integrity Checks Calculate SHA-256 at write and verify at restore to detect bit-rot or tampering. Automated Sandbox Verification Use isolated VMs to automatically boot and verify recovery environments. The Strategic Mandate "A backup is only valid if it restores within the required RTO and meets the required RPO." RECOVERY SUCCESS CRITERIA Measured Restore Time ≤ RTO Key Distinction: Having a 100% "Success" rate in backup logs does not guarantee a successful restoration. Verification must be proactive, periodic, and evidence-based. Warning: An untested DRP is a liability, not a protection. Restoration speed is often bottlenecked by bandwidth, decryption, or media latency. UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Fault Tolerance vs High Availability vs Disaster Recovery Component-Level Failure System/Service Interruption Site-Wide Catastrophe Fault Tolerance PRIMARY OBJECTIVE Zero downtime and continuous service. MECHANISM Hardware redundancy; duplicate components operating in parallel. SCOPE OF LOSS Zero data loss; zero interruption during failure. Highest cost and complexity; handles component failures transparently. High Availability PRIMARY OBJECTIVE Minimized downtime and uptime maximization. MECHANISM Clustering, load balancing, and automated failover to standby. SCOPE OF LOSS Brief interruption; minimal/negligible data loss. Focuses on service continuity despite system software or server failures. Disaster Recovery PRIMARY OBJECTIVE Restoration of service after site loss. MECHANISM Off-site backups, secondary data centers, and DRP runbooks. SCOPE OF LOSS Acceptable downtime (RTO) and data loss (RPO). The final safety net; involves manual or scripted site-level rebuilding. UNIT II: DISASTER RECOVERY & FAULT TOLERANCE Redundancy Layers & Single Points of Failure Power Redundancy Uninterruptible Power Supplies (UPS) Backup Diesel Generators Dual Power Distribution Units (PDUs) Redundant Utility Grid Feeds Hardware Redundancy Redundant Power Supplies (PSUs) NIC Teaming / Link Aggregation ECC Memory & Hot-spare CPU RAID Storage (Local Resilience) Network Redundancy Dual ISPs (Diverse Path Routing) Redundant Core Switches & Routers BGP Multi-homing Load Balancers (Traffic Distribution) Software Redundancy Active-Passive/Active-Active Clusters Microservices / Containers Application Failover Logic Health Check Heartbeats Data Redundancy Database Replication (Async/Sync) Distributed File Systems Cloud Object Versioning Continuous Data Protection (CDP) Geographic Redundancy Multi-Region Deployment Cloud Availability Zones (AZ) Off-site Disaster Recovery Centers Global Load Balancing (GSLB) UNIT II: DISASTER RECOVERY & FAULT TOLERANCE RAID Building Blocks Striping Splits data into "stripes" or chunks across multiple physical disks. Increases I/O performance. Allows parallel read/writes. Increases capacity utilization. Provides zero redundancy. Mirroring Simultaneously writes the exact same data to two or more disks. Full data duplication. Highest fault tolerance. Shortest recovery time. 50% storage overhead. Parity Calculates binary logic (XOR) to store mathematical checksums. Reconstructs lost data. Balances capacity and safety. Efficient space utilization. Write performance penalty. CRUCIAL DISTINCTION: RAID IS NOT A BACKUP RAID protects against physical hardware (disk) failure, but it does NOT provide data protection for other threats. RAID offers zero protection against: Accidental Deletion Ransomware Encryption Data Corruption Physical Theft Site Disaster (Fire/Flood) Teaching Point: If you delete a file on a RAID 1 (Mirror), the deletion is instantly "mirrored" to the other disk. The data is gone. UNIT II: DISASTER RECOVERY & FAULT TOLERANCE RAID Levels: Choosing the Right Design RAID 0: Striping Min Drives: 2 Fault Tolerance: 0 (Zero) Performance: Max Read/Write Best Use: Temporary files, non-critical data. RAID 1: Mirroring Min Drives: 2 Fault Tolerance: 1 Drive Performance: High Read, Normal Write Efficiency: 50% Usable Capacity RAID 5: Distributed Parity Min Drives: 3 Fault Tolerance: 1 Drive Risk: Long rebuild times on large drives. Best Use: Standard storage/file servers. RAID 6: Dual Parity Min Drives: 4 Fault Tolerance: 2 Drives Performance: Slower writes than RAID 5. Advantage: Survives failure during rebuild. RAID 10: 1+0 (Hybrid) Min Drives: 4 Fault Tolerance: 1 per mirror pair. Performance: Peak IOPS/Low latency. Best Use: Databases and high-load apps. The Hot Spare Strategy An idle, powered-on drive ready to automatically replace a failed disk, initiating an immediate background rebuild to minimize the window of vulnerability. CRITICAL REMINDER: RAID is not a backup. It provides hardware availability but cannot recover from accidental deletion, ransomware, or site disaster. UNIT II: DISASTER RECOVERY & FAULT TOLERANCE High-Availability Clustering and Failover Active-Passive Model SERVICE OWNERSHIP Node A ACTIVE Node B Standby Cold/Warm Standby: Secondary node remains idle until the primary fails. Simple Management: No complex state synchronization required for concurrent traffic. Lower Utilization: 50% of hardware resources are typically dormant. Failover Time: Service pause while standby node takes over resources. Active-Active Model SERVICE OWNERSHIP Node A ACTIVE Node B ACTIVE Load Distribution: All nodes actively serve requests simultaneously. High Throughput: Maximizes hardware utilization and capacity. Complexity: Requires robust session state sharing and locking. Resilience: Surviving nodes absorb the load of the failed component. The Failover Mechanism Heartbeat Monitoring A dedicated network link or "keep-alive" signal between nodes. If the heartbeat is lost, the cluster initiates a failure detection protocol. Automatic Failover Process: (1) Detect Failure → (2) Verify (Quorum) → (3) Reassign IP/Services → (4) Update Load Balancer → (5) Restore Traffic.