Getting Started with staRburst

Introduction

staRburst makes it trivial to scale your parallel R code from your laptop to 100+ AWS workers. This vignette walks through setup and common usage patterns.

Which API should I use?

staRburst offers a few entry points. They share the same engine — pick the one that fits how your code is written:

Need Use
Map a function over inputs (simplest) starburst_map()
Existing future / furrr code plan(starburst) then future_map()
Long-running job that survives disconnects starburst_session()
Explicit, reusable worker cluster starburst_cluster()
Configure account/backend defaults starburst_config()

If you’re new, start with starburst_map() — the next section walks through a first job with it. If you already have future/furrr code, skip to Migrating existing future / furrr code; the two styles are equivalent and shown side by side there.

Which backend? EC2 (default) vs Fargate

Every API above runs on one of two compute backends, chosen with launch_type:

  • EC2 — the default and recommended choice. Faster (no cold start), cheaper with Spot (the default), and more flexible (any instance type, warm pools). starburst_setup() provisions its capacity provider for you, so it works out of the box.
  • Fargate — an optional serverless alternative (launch_type = "FARGATE"). Fully managed and needs no starburst_setup_ec2 / capacity provider, but has cold-start latency and is bounded by your account’s Fargate vCPU quota. Reach for it if you specifically want task-based serverless execution.

You don’t have to choose up front — leave the default (EC2) and add launch_type = "FARGATE" later if you want to compare.

How it works (the mental model)

        Local R session
              |
              |  (1) serialize your function + inputs + detected globals (qs2)
              |      AWS credentials are read HERE, from your local environment
              v
   S3 bucket  +  staRburst control plane (ECS/ECR)
              |
              |  (2) launch workers: EC2 (default, Spot) or Fargate
              |      each worker's R packages come from your renv.lock, baked
              |      into a Docker image cached in ECR
              v
   Remote R workers  (one task per input element)
              |
              |  (3) each worker pulls a task from S3, runs it, writes the
              |      result + logs back to S3
              v
   Local R session  <-- results collected in order (starburst_map / cluster)
        or
   Detached-session store in S3  <-- collected later via starburst_session_attach()

What to know from this picture:

Installation

# Install from GitHub
remotes::install_github("scttfrdmn/starburst")

One-Time Setup

Before using staRburst, you need to configure AWS resources. This only needs to be done once.

library(starburst)

# Interactive setup wizard
starburst_setup()

This will: - Validate your AWS credentials - Create an S3 bucket for data transfer - Create an ECR repository for Docker images - Set up an ECS cluster and VPC resources - Provision the default EC2 capacity (launch template + Auto Scaling Group + capacity provider for c7g.xlarge) so the default EC2 backend works immediately — created at zero instances, so no compute cost until you run a job (skip with setup_ec2 = FALSE if you only use Fargate) - Check Fargate quotas and offer to request increases - Build the initial worker image

Provisioning the AWS resources takes about 2 minutes. On the first run, staRburst also builds the worker Docker image, which adds 5–10 minutes (it is cached and reused afterwards, and you can skip it with starburst_setup(build_image = FALSE) and build lazily on first job launch).

Your first job

The simplest way to use staRburst is starburst_map() — hand it your inputs and a function, and it runs the function over each input on AWS workers:

library(starburst)

# Define your work
expensive_simulation <- function(i) {
  # Some computation that takes a few minutes
  results <- replicate(1000, {
    x <- rnorm(10000)
    mean(x^2)
  })
  mean(results)
}

# Run across 50 cloud workers (EC2 by default)
results <- starburst_map(1:100, expensive_simulation, workers = 50)
#> [Starting] Starting starburst cluster with 50 workers
#> [Status] Processing 100 items with 50 workers
#> [Starting] Submitting 100 tasks...
#> [Wait] Progress: 100/100 (128.0s)
#> [OK] Completed in 128.0 seconds
#> [Cost] Estimated cost: $0.85

That is the whole workflow: starburst_setup() once, then starburst_map() per job. Everything below builds on this.

Migrating existing future / furrr code

If you already use future or furrr, you don’t need starburst_map() — set the plan to starburst and your existing future_map() / future_lapply() calls run on AWS unchanged. The two styles are equivalent; use whichever matches your code:

library(furrr)
library(starburst)

# Local baseline
plan(sequential)
results_local <- future_map(1:100, expensive_simulation)

# Same code, now on 50 cloud workers — just change the plan
plan(starburst, workers = 50)
results_cloud <- future_map(1:100, expensive_simulation)

# Results are identical to the local run
identical(results_local, results_cloud)
#> [1] TRUE

The two calls below do the same thing — pick the one that fits your codebase:

# Direct API
results <- starburst_map(1:100, expensive_simulation, workers = 50)

# future / furrr
plan(starburst, workers = 50)
results <- future_map(1:100, expensive_simulation)

Real-world examples

Rather than repeat toy snippets here, the article gallery carries complete, measured end-to-end examples — each with real runtimes, costs, and the batching choices that make (or break) a workload:

Before you size a job, read the two guides that report real numbers on when the cloud wins, when it loses, and how to pick task/worker counts: Workload Shapes and Performance. The short version: each task should be real work (seconds or more), and thousands of tiny tasks should be batched into dozens–hundreds — see Batch Small Tasks below.

Working with Data

Base R’s read.csv()/readRDS() cannot read s3:// URLs — you need an S3 client on the worker. A small helper keeps the examples below concrete and copy-pasteable (the worker image already includes paws.storage):

# Download an s3://bucket/key object to a temp file and read it on the worker.
read_s3 <- function(s3_uri, reader = readRDS) {
  m <- regmatches(s3_uri, regexec("^s3://([^/]+)/(.+)$", s3_uri))[[1]]
  bucket <- m[2]; key <- m[3]
  tmp <- tempfile()
  obj <- paws.storage::s3()$get_object(Bucket = bucket, Key = key)
  writeBin(obj$Body, tmp)
  reader(tmp)
}

Data Already in S3

If your data is already in S3, workers can read it directly:

plan(starburst, workers = 50)

results <- future_map(file_list, function(file) {
  # Each worker pulls its file from S3 (read_s3 defined above)
  data <- read_s3(sprintf("s3://my-bucket/%s", file), reader = read.csv)
  process(data)
})

Uploading Local Data

For smaller datasets, you can pass data as arguments:

# Load data locally
data <- read.csv("local_file.csv")

# staRburst automatically uploads to S3 and distributes
plan(starburst, workers = 50)

# `replicates` is your vector of inputs (e.g. bootstrap replicate ids)
results <- future_map(replicates, function(i) {
  # Each worker gets a copy of 'data'
  bootstrap_analysis(data, i)
})

Large Data Optimization

For very large objects, upload once to S3 yourself and have each worker read it from there, rather than serializing the object into every task:

# Upload once from your machine
paws.storage::s3()$put_object(
  Bucket = "my-bucket", Key = "large_data.rds",
  Body = "huge_file.rds"
)
s3_path <- "s3://my-bucket/large_data.rds"

# Workers read from S3 inside the task (read_s3 helper defined above)
plan(starburst, workers = 100)

# `tasks` is your vector of work items
results <- future_map(tasks, function(i) {
  data <- read_s3(s3_path)   # readRDS by default
  process(data, i)
})

Cost Management

Estimate Costs

# Check cost before running
plan(starburst, workers = 100, cpu = 4, memory = "8GB")
#> Estimated cost: ~$3.50/hour

Set Cost Limits

# Set an hourly cost ceiling (USD/hour) — jobs above this rate won't start
starburst_config(
  max_hourly_cost = 10,       # Don't start jobs estimated over $10/hour
  cost_alert_threshold = 5     # Warn at $5/hour
)

# Now jobs exceeding limit will error before starting
plan(starburst, workers = 1000)  # Would cost ~$35/hour
#> Error: Estimated cost ($35/hr) exceeds limit ($10/hr)

Cost estimates

After a run, staRburst reports an estimated cost — the measured worker runtime multiplied by the current AWS price for your instance type (On-Demand or Spot), looked up live from the AWS Pricing API (and cached; it falls back to built-in rates offline). It is a close estimate, not a figure from AWS billing / Cost Explorer.

plan(starburst, workers = 50)

results <- future_map(data, process)

#> Cluster runtime: ~23 minutes
#> Estimated cost: ~$1.34

Quota Management (Fargate backend)

This section applies to the Fargate backend. On the default EC2 backend, worker count is bounded by your normal EC2 On-Demand / Spot instance limits, not the Fargate vCPU quota below. If you stay on EC2 you can usually skip this.

Check Your Quota

starburst_quota_status()
#> Fargate vCPU Quota: 100 / 100 used
#> Allows: ~25 workers with 4 vCPUs each
#>
#> Recommended: Request increase to 500 vCPUs

Request Quota Increase

starburst_request_quota_increase(vcpus = 500)
#> Requesting Fargate vCPU quota increase:
#>   Current: 100 vCPUs
#>   Requested: 500 vCPUs
#>
#> [OK] Quota increase requested (Case ID: 12345678)
#> [OK] AWS typically approves within 1-24 hours

Wave-Based Execution

If you request more workers than your quota allows, staRburst automatically uses wave-based execution:

# Quota allows 25 workers, but you request 100
plan(starburst, workers = 100, cpu = 4)

#> [!] Requested: 100 workers (400 vCPUs)
#> [!] Current quota: 100 vCPUs (allows 25 workers max)
#>
#> [Plan] Execution plan:
#>   - Running in 4 waves of 25 workers each
#>
#> [TIP] Request quota increase to 500 vCPUs? [y/n]: y
#>
#> [OK] Quota increase requested
#> [Starting] Starting wave 1 (25 workers)...

results <- future_map(inputs, expensive_function)

#> [Wave] Wave 1: 100% complete (250 tasks)
#> [Wave] Wave 2: 100% complete (500 tasks)
#> [Wave] Wave 3: 100% complete (750 tasks)
#> [Wave] Wave 4: 100% complete (1000 tasks)

Troubleshooting

View Worker Logs

# View logs from most recent cluster
starburst_logs()

# View logs from specific task
starburst_logs(task_id = "abc-123")

# View last 100 log lines
starburst_logs(last_n = 100)

Check Cluster Status

starburst_status()
#> Active Clusters:
#>   • starburst-xyz123: 50 workers running
#>   • starburst-abc456: 25 workers running

Common Issues

Environment mismatch: Packages not found on workers

# Rebuild environment
starburst_rebuild_environment()

Task failures: Some tasks failing

# Check logs
starburst_logs(task_id = "failed-task-id")

# Often due to memory limits - increase worker memory
plan(starburst, workers = 50, memory = "16GB")  # Default is 8GB

Slow data transfer: Large objects taking too long

# Use Arrow for data frames
library(arrow)
write_parquet(my_data, "s3://bucket/data.parquet")

# Workers read Arrow
results <- future_map(1:100, function(i) {
  data <- read_parquet("s3://bucket/data.parquet")
  process(data, i)
})

Best Practices

1. Use for Right-Sized Workloads

Good: Each task takes >5 minutes

# 100 tasks, each takes 10 minutes
# Local: 1000 minutes, Cloud: ~10 minutes

Bad: Each task takes <1 minute

# 10000 tasks, each takes 30 seconds
# Startup overhead (45s) dominates

2. Batch Small Tasks

Instead of:

# 10,000 tiny tasks
results <- future_map(1:10000, small_function)

Do:

# 100 batches of 100 tasks each
batches <- split(1:10000, ceiling(seq_along(1:10000) / 100))

results <- future_map(batches, function(batch) {
  lapply(batch, small_function)
})

# Flatten results
results <- unlist(results, recursive = FALSE)

3. Use S3 for Large Data

Don’t:

big_data <- read.csv("10GB_file.csv")  # Upload for every task
results <- future_map(1:1000, function(i) process(big_data, i))

Do:

# Upload once to S3
tmp <- tempfile(fileext = ".csv"); write.csv(big_data, tmp, row.names = FALSE)
paws.storage::s3()$put_object(Bucket = "bucket", Key = "big_data.csv", Body = tmp)

# Workers read from S3 (read_s3 helper from "Working with Data" above)
results <- future_map(tasks, function(i) {
  data <- read_s3("s3://bucket/big_data.csv", reader = read.csv)
  process(data, i)
})

4. Set Reasonable Limits

starburst_config(
  max_hourly_cost = 50,            # Cap the hourly rate (USD/hour)
  cost_alert_threshold = 25        # Get warned early
)

5. Clean Up

# staRburst auto-cleans, but you can force it
plan(sequential)  # Switch back to local
# Old cluster resources are cleaned up automatically

Advanced: Custom Configuration

CPU and Memory

# High CPU, low memory (CPU-bound work)
plan(starburst, workers = 50, cpu = 8, memory = "16GB")

# Low CPU, high memory (memory-bound work)
plan(starburst, workers = 25, cpu = 4, memory = "32GB")

Timeout

# Increase timeout for long-running tasks (default 1 hour)
plan(starburst, workers = 10, timeout = 7200)  # 2 hours

Region

# Use specific region (default from config)
plan(starburst, workers = 50, region = "us-west-2")

Next Steps

Getting Help