11: Data Management in the Cloud
- Page ID
- 128104
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)Cloud computing has transformed how organizations store, manage, and analyze data. In traditional on-premises settings, data management focused on centralized databases and direct-attached storage. In contrast, cloud data management emphasizes distributed systems, scalability, and resiliency to meet the demands of modern applications. This chapter explores the principles and practices of managing data in cloud environments, blending academic concepts with practical applications relevant to cloud project management. We will cover fundamental cloud data management principles, examine different cloud storage models (block, object, and file storage) and distributed file systems, compare SQL, NoSQL, and NewSQL databases, discuss data replication strategies and consistency models, outline backup and recovery approaches in the cloud, and delve into name services and metadata management in distributed systems. Throughout, real-world examples (such as Netflix’s cloud architecture and Google’s Spanner database) and cloud provider services (AWS, Azure, GCP) are included to illustrate these concepts in practice.
Learning Objectives
After completing this chapter, students will be able to:
- Explain cloud data management principles: durability, availability, scalability, consistency, security, and cost efficiency.
- Compare cloud storage models: object storage, file storage, and block storage, including their trade-offs.
- Describe distributed file systems (HDFS, CephFS) and their role in big data and analytics.
- Differentiate between SQL, NoSQL, and NewSQL databases and select appropriate types for given workloads.
- Analyze data replication strategies and consistency models (strong, eventual, causal) in distributed systems.
- Design backup and disaster recovery strategies for cloud-based data.
- Evaluate real-world cloud data architectures including Netflix and Google Spanner.
- 11.1: Cloud Data Management Principles
- Cloud data management focuses on durability and availability (replication across servers/sites), scalability (horizontal growth), cost efficiency (tiering and pay-as-you-go), consistency and integrity in distributed systems, and security and compliance (encryption, IAM, governance).
- 11.2: Cloud Storage Modes and Distributed File Systems
- Cloud storage offers block (disks for VMs), object (scalable, flat namespace for data lakes/media), and file (hierarchical, shared access for legacy apps) models, with distributed file systems (e.g., HDFS) extending file storage to cluster scale via metadata services, all chosen based on access patterns, performance, and cost tradeoffs.
- 11.3: SQL, NoSQL, and NewSQL Databases - Definitions, Use Cases, and Comparisons
- Relational databases (SQL) provide tables, schemas, ACID, and SQL via managed services (e.g., RDS, Azure SQL, Cloud SQL, Aurora) for transactional and structured workloads; NoSQL (key-value, document, wide-column, graph) offers flexible schema and horizontal scale with often eventual consistency (e.g., DynamoDB, Cosmos DB, Bigtable) for high throughput; NewSQL (e.g., Spanner, CockroachDB, Aurora Global) delivers distributed SQL with ACID for global scale and strong consistency—with SQL favored f
- 11.4: Data Replication Strategies and Consistency Models
- Replication strategies—single-leader, multi-leader (with conflict handling), quorum-based (tunable read/write), and geo-replication—combined with sync vs. async replication, shape durability and latency; consistency ranges from strong (linearizability) through sequential, causal, and eventual, with CAP framing the trade-off between consistency and availability under partitions, and many cloud databases (e.g., Cosmos DB, DynamoDB) offering tunable consistency to match workload needs.
- 11.5: Backup and Recovery Strategies in the Cloud
- Backup and recovery in the cloud use snapshots, point-in-time restore, and cross-region copies; follow a 3-2-1 style (multiple copies, media, off-site); use automation (e.g., AWS Backup, Azure Backup) and lifecycle policies (e.g., Glacier); define RTO/RPO and test restores; and protect backup data (encryption, access control, immutability). Replication is not backup—backups must protect against logical errors and deletion.
- 11.6: Name Services and Metadata Management in Distributed Environments
- DNS and service discovery (e.g., Eureka, Consul, Cloud Map, Kubernetes DNS) map logical names to endpoints so components find each other as instances scale and change. Metadata Management: In DFS and storage, metadata services (e.g., HDFS NameNode, Ceph MDS) hold namespace and block location; object stores avoid a single metadata server by partitioning the key space. Design for HA and scalability of naming and metadata to avoid single points of failure.
- 11.7: Conclusion and Key Takeaways
- The conclusion stresses choosing storage and database types to match workload needs, planning replication and consistency (including CAP trade-offs), implementing backup and DR with automation and testing, designing naming and metadata for HA and scale, and using managed cloud services while understanding the underlying concepts so data remains durable, scalable, and manageable in the cloud.


