Infrastructure Stability & Scale Readiness Audit

Sholder
Sholder

Other Engineering · Contractor

Posted on Sep 14, 2026
Request for Proposal · Fixed-scope contract · Remote · Firm or individual

Infrastructure Stability & Scale Readiness Audit

Tell us whether our infrastructure holds at the scale ahead — and exactly what to change first if it doesn't.

The short version

Sholder is a live platform that connects people with a personal network of trusted support — live video sessions, messaging, and an AI layer that makes every relationship more intentional. Our product is in production on AWS and our members are real. We're a small team, and most of our engineering is executed by an AI development system working inside a codebase and infrastructure built to move quickly without losing quality.

We're about to open the platform to a materially larger audience. Before we do, we want an independent, expert answer to one question:

Will the infrastructure we have today hold at the scale we're about to ask of it — and if not, exactly what do we change first?

This is not a security review. That's covered separately. This engagement is about stability and scale: does it stay up, does it stay fast, does it recover, and does it grow without surprises.

What you'd be looking at

At a high level here, and in full detail under NDA after selection:

  • A single-region AWS estate defined in Terraform: containerized services on ECS Fargate behind an application load balancer, a managed Multi-AZ Postgres database, a static single-page application on S3 and CloudFront, and Cloudflare at the edge.
  • A Go API and a Node real-time messaging service, plus a React web application.
  • Background work that runs inside the API process rather than in a separate worker tier.
  • Real-time dependencies: a third-party video SDK for live sessions, a WebSocket service for chat and guided AI conversations, an LLM API, an SMS provider, a payments provider, and a payroll-partner integration.
  • CI/CD through GitHub Actions with migration gating and post-deploy smoke tests.
  • Existing monitoring, alerting, backup, and recovery arrangements.

Scope

In scope

  1. Architecture review for single points of failure, capacity ceilings, and failure modes under load — across compute, data, networking, edge, and third-party dependencies.
  2. Horizontal scaling readiness: which services can safely run more than one instance today, which cannot, and what each one needs.
  3. Database posture: instance sizing, connection management, growth headroom, migration safety, and recovery objectives.
  4. Real-time path: WebSocket and live-video session capacity and behavior during deploys, instance loss, and partner outages.
  5. Observability and operations: whether alerting, dashboards, runbooks, and on-call arrangements are sufficient to detect and respond at the target scale.
  6. Deploy and rollback safety at higher traffic.
  7. Resilience: backup, restore, and regional-failure posture, measured against recovery objectives we'll define together in week one.
  8. A load or soak test against a non-production environment is welcome if you consider it necessary to reach a verdict. Tell us in your proposal whether you'll do one and what you need from us.

Out of scope

  • Security testing, penetration testing, and compliance work.
  • Application code quality, product behavior, and UX.
  • Cost optimization, except where cost directly affects stability or scale.
  • Our internal developer tooling and marketing website.
  • Performing the remediation. We'll execute it ourselves.

Deliverables

  1. A readiness verdict against a written scale envelope — Ready, Ready with conditions, or Not ready — with the reasoning.
  2. A findings report: each finding with severity, the failure scenario it produces, and evidence.
  3. A prioritized remediation roadmap, ordered by risk reduction per unit of effort, with rough effort estimates.
  4. A live readout with our leadership and engineering.

One requirement that's unusual

Our remediation will be executed by our AI development system. Every recommendation must be written as a specific, verifiable acceptance criterion, not a heading. "Add autoscaling" is not actionable for us. "Service X should scale between N and M tasks on metric Y at threshold Z, verified by test T" is.

What we provide

  • Read-only access to the AWS accounts, the infrastructure-as-code repository, architecture documentation, and existing runbooks.
  • Read access to the application repositories where needed to trace a concern.
  • Up to two hours per week of interview time with the team, plus async questions with same-day answers.
  • A written scale envelope, finalized with you in the first week: target concurrent live sessions, active members, and traffic multiple over the next 12 months.

All of it under a mutual NDA, signed before access is granted.

Timeline

Milestone Date
Questions accepted Sep 8 – 12
Proposals due Sep 14
Selection and NDA Sep 17
Kickoff and access Sep 19
Interim findings Sep 28
Final report and readout Oct 7

How we'll evaluate

  • Depth of relevant experience with systems shaped like ours.
  • Quality and precision of the sample report.
  • Realism of the plan against the 30-day window.
  • Fit with our acceptance-criteria requirement.
  • Price — last.

Engagement

Fee
Propose your fee
fixed-fee preferred
Window
30 days
proposals due Sep 14 · readout Oct 7

Fixed-fee proposals preferred. Tell us your price, payment terms, and assumptions. Price is the last thing we evaluate, not the first.

To apply

Keep it short. We read everything, and we prefer a tight proposal to a long one. Send yours to taylor@sholder.com with the subject line "Infra audit RFP", covering:

  1. Your approach, in your own words, and how you would reach a verdict inside three weeks.
  2. Who does the work, with relevant experience on AWS ECS, Postgres, and real-time systems at the scale in question.
  3. Two references for comparable engagements.
  4. A redacted sample of a prior audit report, so we can see how you write findings.
  5. Fixed price, payment terms, and any assumptions.
  6. What you need from us that is not listed here.

One more thing: include the first question you would ask us. It's the fastest signal we have for how you think.

Send a proposal

We read everything.