Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Twitter did not move its entire service or all of its data centers to Google Cloud. Beginning in May 2018, it moved selected data-platform workloads—especially cold storage and ad hoc Hadoop compute—while keeping real-time and production Hadoop clusters on premises. The project later grew into a broader analytics modernization using Google Cloud Storage, BigQuery, Dataflow, Bigtable and managed Hadoop services.
What Twitter announced in 2018
On May 3, 2018, Twitter announced a collaboration with Google Cloud focused on two categories: cold data storage and flexible-compute Hadoop clusters. Twitter said its Hadoop file systems contained more than 300 PB of data spread across tens of thousands of servers. The announcement did not say that user-facing traffic, core application systems or every data-center workload would move to Google Cloud.
The distinction matters because Twitter operated several Hadoop environments with different risk and latency profiles. Cold data and ad hoc analysis could tolerate more elastic, batch-oriented infrastructure. Real-time ingestion, production processing and serving paths had tighter operational requirements.
Twitter cited faster capacity provisioning, infrastructure flexibility, access to managed tools, security improvements and better disaster-recovery options as reasons for working with Google Cloud. The 2018 announcement describes those goals, but does not establish a completed 300-PB transfer at that date.
#1 Best Overall
How large was the starting environment?
The more-than-300-PB figure described Twitter’s Hadoop file systems, not the amount instantly relocated in 2018. A contemporaneous industry report said Hadoop represented close to 20% of Twitter’s hardware in a January 2017 snapshot. That was the broader Hadoop footprint; it does not mean Twitter moved 20% of its infrastructure to Google Cloud.
Because the estate comprised multiple clusters, migration had to be planned by workload rather than as a single server-replacement exercise. Storage density, CPU utilization, latency, data sensitivity and operational dependencies differed from one cluster to another.
The hybrid architecture Twitter chose
Twitter investigated an all-in cloud move but described it as too substantial to undertake at that stage. Instead, it selected workloads where cloud elasticity and managed services offered benefits without immediately moving the most operationally sensitive systems. Its 2019 architecture description gives this split:
| Workload | Location in the described architecture |
|---|---|
| Real-time ingestion and serving | Twitter data centers |
| Production Hadoop processing | Twitter data centers |
| Ad hoc analysis | Google Cloud |
| Cold storage | Google Cloud |
| Shared object-storage layer | Google Cloud Storage |
| Managed analytics and processing | BigQuery, Dataflow and related services |
In other words, this was a selective, hybrid migration—not an abandonment of Twitter’s data centers. The architecture article identifies the Ad Hoc and Cold Storage clusters as moving to the cloud while Real Time and Production Hadoop clusters remained on premises: Twitter’s 2019 architecture account.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Why cold storage and ad hoc Hadoop were early candidates
Cold does not mean unimportant
“Cold” generally means data is accessed less frequently or is less latency-sensitive. Historical data can still be essential for machine learning, investigations, recovery, compliance and long-term analysis. Its access pattern, rather than its business value, made it more suitable for object storage and elastic processing.
Rank #2
Twitter described its cold clusters as storage-dense but carrying underused CPU. Separating storage from compute allowed data to persist without keeping a large amount of idle processing capacity attached to it.
Ad hoc workloads are difficult to size
Ad hoc analysis can arrive in bursts. A cloud environment can provision temporary compute for a job instead of requiring a permanently oversized on-premises cluster. This is useful when data volumes are large but demand varies significantly over time.
Real-time systems carry different risks
Real-time ingestion and production processing depend on predictable latency, tightly controlled failure behavior and many application-specific assumptions. Keeping those clusters on premises while proving the cloud model on less latency-sensitive workloads reduced the blast radius of early migration problems.
Separating compute from storage
Traditional Hadoop deployments commonly colocate HDFS storage and compute on the same servers. Cloud object storage changes that design. Data can remain in Cloud Storage while different compute environments are created for different jobs.
- Compute capacity can be added without purchasing more disks.
- Storage can grow without adding equivalent CPU capacity.
- Temporary clusters can process persistent data and then be removed.
- Different machine types can be matched to different jobs.
- Managed services can handle parts of cluster, stream-processing or warehouse operations.
Google describes Dataproc as a managed Hadoop and Spark service that can process data in Cloud Storage and write results to Cloud Storage, BigQuery or Bigtable: Google’s Dataproc migration guidance. Disaggregation can improve utilization and flexibility, but it is not an automatic cost reduction. Network traffic, request charges, storage classes, duplicated data and inefficient jobs can outweigh hardware savings.
Rank #3
How Twitter approached more than 300 PB
A petabyte-scale migration is not a one-time “copy HDFS to object storage” operation. Twitter’s data continued to change while the transfer ran: new partitions were produced, older data aged out and some content required sensitive-data scrubbing.
- Initial bulk replication: copy the existing eligible data set into Google Cloud Storage.
- Continuous synchronization: keep new, changed and deleted partitions aligned while both environments operate.
- Data cleansing: apply required privacy or sensitive-data filtering before cloud copies become usable.
- Validation and reconciliation: compare manifests, partition counts, checksums or equivalent controls and investigate gaps.
- Workload cutover: move selected jobs only after reads, writes, permissions and performance have been verified.
- Hybrid operation and rollback: retain a controlled dual-running path while failures and replay behavior are understood.
Twitter said it aimed to replicate more than 300 PB into Cloud Storage and used a continuously synchronized process rather than treating the data as a static archive. The company also said physical-device transfer was not practical for the live, continually changing data set described in its architecture account.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsConnectivity was a design constraint
Twitter’s 2019 description said private access to Google Cloud initially used Dedicated Interconnect, which was limited to 10 Gbps in the setup and period described. That number is not a universal Google Cloud limit; it characterizes the connectivity available to Twitter at that stage: the architecture post.
In a later ingestion design, Twitter described an approximately 77 Gbit/s proxy bottleneck while moving data from on-premises Hadoop into Google Cloud. Again, this was a limitation of that particular path, not a general ceiling for every Google Cloud transfer: the 2022 retrospective.
At these scales, bandwidth planning must include ongoing synchronization, retries, validation traffic and the reads generated by active analytics—not just the initial copy window.
Rank #4
HDFS and object storage are not interchangeable
Applications written for HDFS often assume filesystem-like behavior. Object storage has different semantics and performance characteristics. Migration teams need to test:
- Rename and delete behavior, especially where jobs expect atomic filesystem operations.
- Directory and file-listing behavior and the cost of large metadata operations.
- Small-file accumulation, which can make planning and query execution inefficient.
- Permissions, service identities, encryption and audit logging.
- Retry and replay behavior so failed transfers do not create silent duplicates.
- Partitioning and file layout for downstream analytics.
Twitter and Google also worked on the Cloud Storage Connector for Hadoop. Google reported SQL testing on a dataset larger than 20 PB and described performance work for Parquet and ORC, including selective reads and cooperative locking: Google’s connector release account. Columnar formats, predicate pushdown and range reads can reduce unnecessary scanning, but they do not eliminate the need for sound data layout and metadata management.
From Hadoop migration to a broader analytics platform
The cloud project expanded beyond storage. Twitter used Cloud Storage as a shared layer for services including Cloud Dataflow, Cloud Dataproc and BigQuery. Its BigQuery rollout included a company-wide alpha with Data Studio in November 2018, followed by general availability at Twitter in April 2021.
Twitter’s later platform work described Cloud Replicator and Airflow-based workflows for moving and managing data. BigQuery provided centralized SQL analytics, while Dataflow handled managed batch and streaming transformations. Bigtable served low-latency access patterns in some analytics and advertising systems. Google’s account of the advertising platform describes the combination of BigQuery, Bigtable, Cloud Storage and Dataflow: the advertising analytics case study.
The scale increased over time. An earlier BigQuery phase reported approximately 8,000 queries processing nearly 100 PB in one month. By 2022, Twitter said employees were running more than 10 million queries per month on almost one exabyte of data. Those figures describe different stages and scopes; the 2022 analytics environment should not be presented as the amount physically moved in the original 2018 project. The later retrospective is available at Twitter’s exabyte-scale data-access article.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Best Value
What the migration delivered—and what it did not prove
Documented benefits
- Faster capacity provisioning than acquiring and installing equivalent data-center hardware.
- More flexible separation of persistent data and temporary compute.
- Access to managed Hadoop, streaming, warehouse and serving services.
- Additional replication and disaster-recovery options, subject to each workload’s recovery design.
- Broader analyst access through governed BigQuery projects rather than direct dependence on low-level Hadoop tools.
Twitter’s BigQuery account emphasized governance, resource allocation and easier enterprise-wide data use: the BigQuery rollout article. Twitter also used flat-rate slots in that period to make warehouse spending more predictable than relying solely on on-demand query pricing.
Trade-offs and failure modes
- Transfer and query costs: repeated cross-region or cross-service reads can become expensive.
- Semantic mismatch: HDFS assumptions about renames, locality or mutable files may require code changes.
- Governance drift: permissions, retention rules and sensitive-data deletion must remain consistent across environments.
- Hybrid complexity: operators must monitor two environments, synchronize data and coordinate incident response.
- Lock-in: heavy dependence on BigQuery, Dataflow, Bigtable and Google-specific IAM or APIs can make later migration harder.
- Unpredictable scans: poorly partitioned BigQuery tables and uncontrolled ad hoc queries can undermine cost planning.
- Premature migration: moving real-time or production clusters before proving performance and rollback procedures increases operational risk.
Storage and compute locations must be designed together. Google announced that certain BigQuery requests reading from multi-region Cloud Storage could incur multi-region data-transfer charges beginning February 1, 2026: Google’s billing notice.
What other large enterprises can learn
- Segment workloads first. Classify data and jobs by latency, statefulness, access frequency, sensitivity and operational criticality.
- Start with reversible candidates. Cold and ad hoc workloads can provide a lower-risk proving ground than real-time serving.
- Treat movement as a service. Build synchronization, deletion handling, validation, monitoring and replay into the architecture.
- Test object-storage behavior early. Measure listing, rename, small-file, locking and metadata behavior with representative jobs.
- Place data and compute deliberately. Region, network path, partitioning and query design affect both latency and cost.
- Establish governance before democratization. Broad SQL access requires project boundaries, service-account controls, auditability and chargeback.
- Keep a rollback path. Dual-running and tested reconciliation are safer than an irreversible cutover.
- Modernize workflows, not just locations. The largest gains may come from managed processing and accessible analytics rather than simply renting replacement servers.
What this says about Twitter’s strategy
The public record describes a staged hybrid-cloud transformation of Twitter’s data platform. It began with cold storage and flexible Hadoop compute, retained real-time and production clusters on premises in the documented 2019 architecture, and later expanded into BigQuery-centered warehousing, Dataflow pipelines and Bigtable-backed serving.
Because the cited engineering posts describe Twitter-era systems, they do not establish that the current X platform in 2026 still uses exactly the same architecture or that every workload ever reached a final cloud-only state.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




