Cloud HPC Cluster - MLOps @ Meta (Facebook)

Start Your Project

Cloud HPC cluster with GPUs for large-scale ML research

Case Study

Our solution

The AI/ML explosion changed the economics of research infrastructure overnight. Meta (formerly Facebook) needed to scale GPU capacity fast — more researchers, more compute, more experiments running in parallel — but on-premises HPC clusters couldn't keep up. On-premises infrastructure has real advantages: customization, performance, security, and control. But it also carries serious disadvantages: massive upfront investment, long ROI cycles, years-long build timelines, and hardware that becomes obsolete before it's fully utilized. When AI research demands accelerate, waiting years to expand capacity is not an option. Cloud HPC offered a different trade-off: more flexibility, faster provisioning, and cost efficiency for overflow capacity and cutting-edge hardware testing. The challenge was that cloud HPC offerings have limited native features and are notoriously difficult to adopt at enterprise scale. Integration with internal services was a hard requirement — not a nice-to-have.

Solution Delivered

Renaiss designed and deployed a production-ready HPC Slurm cluster on AWS using AWS ParallelCluster as the base layer, heavily customized to meet enterprise requirements.
The architecture went well beyond a standard ParallelCluster deployment. Key capabilities built on top of the base layer included secure access for internal users, Unix user management, two-factor authentication, S3 data pipelines, FSx for Lustre support across multiple configurations, Slurm partitions and limits, Slurm accounting, hardware observability, hardware testing frameworks, login nodes, and multi-tenant support across different AWS accounts. Persistent $HOME directories, Lustre eviction policies, and capacity planning tools were also implemented to support research workflows at scale.

Custom safeguards were built specifically for AWS services to prevent runaway costs and enforce governance. Over time, an Azure cluster was added to the stack using Cycle Cloud, expanding the solution to a true multi-cloud environment.Full tech stack: Terraform, Packer, AWS (EC2, EFA, FSx, EFS, S3, SES, SNS, SQS, Step Functions, Cognito, DynamoDB, CloudWatch), PyTorch, NCCL, DUO.

Project Results

  • The platform scaled to support over 500 researchers across more than 20 clusters, spanning 5+ accounts and tenants. At peak, the infrastructure managed more than 6,000 GPUs under active management and multiple petabytes of data on S3 and FSx.
  • The engagement had an impact beyond the client: AWS ParallelCluster incorporated several ideas developed during this project into its roadmap, a recognition of the technical depth and novelty of the work Renaiss contributed.
  • The result was a research infrastructure that could scale with AI demand — provisioning new clusters in hours instead of years, supporting hundreds of researchers simultaneously, and integrating seamlessly with internal enterprise systems.

+500 researchers

+500 researchers

A research team at full speed, without the friction that on-prem environments create.

+20 clusters

+20 clusters

+Twenty environments running in parallel, each tuned to its team and workload.

+5 accounts/tenants

+5 accounts/tenants

Isolated environments per team, with centralized governance across every account.

+5000 GPUs under management

+5000 GPUs under management

Five thousand GPUs orchestrated, scheduled, and billed with production-grade precision.

multiple PB on S3/FX

multiple PB on S3/FX

Five thousand GPUs orchestrated, scheduled, and billed with production-grade precision.

AWS ParallelCluster took many ideas from this engagement

AWS ParallelCluster took many ideas from this engagement

The work shaped AWS ParallelCluster's roadmap. A signal the architecture was ahead of the product.

From On-Prem Constraints to 5,000 GPUs in the Cloud

Assessment & Architecture Design

We mapped the client's research workflows, data volumes, and GPU demand to define a cloud HPC architecture that could scale without sacrificing security or control.

Infrastructure Deployment

We mapped the client's research workflows, data volumes, and GPU demand to define a cloud HPC architecture that could scale without sacrificing security or control.

Security & Access Configuration

We mapped the client's research workflows, data volumes, and GPU demand to define a cloud HPC architecture that could scale without sacrificing security or control.

Storage & Data Pipeline Integration

We connected multiple FSx for Lustre file systems and S3 pipelines to support petabyte-scale datasets, with automated eviction policies to keep costs under control.

Observability & Ongoing Operations

We set up Slurm accounting, CloudWatch monitoring, and capacity planning tools so the client's team could run autonomously — with full visibility into usage, cost, and hardware performance.

AWS ParallelCluster took many ideas from this engagement

The work shaped AWS ParallelCluster's roadmap. A signal the architecture was ahead of the product.

What is nearshore software development?

Nearshore means hiring a tech team in a country geographically and culturally close to yours — typically within 1–3 time zones. For US companies, that means Latin America. You get real-time collaboration, overlapping work hours, and engineers who operate in English, without the communication friction that comes with offshore teams 10+ hours away.

What time zone does Renaiss operate in?

We're based in Argentina — UTC-3. That means 4–5 hours ahead of the US West Coast and 2 hours ahead of the East Coast. In practice, we maintain a daily overlap of 4–6 hours with most US-based teams, which covers standups, code reviews, and real-time problem-solving without anyone working at midnight.

What cloud services does Renaiss specialize in?

Our work sits at the intersection of cloud infrastructure, application architecture, and AI. Concretely: cloud-native architecture design, infrastructure automation, platform engineering, data pipelines, and GenAI integration. We don't do generic cloud support — we build and run systems that need to scale.

Do you work with AWS, Azure, or GCP?

Our primary depth is in AWS. We also work with Azure and GCP depending on the client's existing stack — the goal is always to work within your environment, not to push a preferred vendor.

Can Renaiss help us modernize a legacy application?

Yes — and it's one of the problems we work on most. App modernization usually means one or more of the following: breaking a monolith into services, re-platforming to cloud-native infrastructure, replacing outdated dependencies, or improving the CI/CD pipeline so your team can ship faster. We start with a technical assessment before recommending any approach.