calyr.aí PARVOTEC

Threadripper 24c + RTX 5090 32GB + 128GB ECC + 10GbE
Local Compute Bridge: MacBook M5 ↔ Workstation ↔ HPC | Training, BO, Simulations

Scientific Node — Compute Bridge Architecture

Design Philosophy

NOT a gaming PC. NOT a replace-HPC machine. A bridge.

LayerHardwarePurpose
Mobile FrontendMacBook M5 Pro (Aorta)Interactive coding, field research, remote control
Local ComputeScientific Node (Parvotec)Training, BO, simulations, overnight batches
Heavy ComputeHPC / VSC ViennaLarge-scale CFD, ensemble methods, production runs

Why Threadripper instead of Intel/AMD EPYC?

Why RTX 5090 not RTX PRO 6000?

RTX 5090 (32 GB)~€4.500–6.000CUDA 12.5, Tensor cores, fast; good for training + BO
RTX PRO 6000 (96 GB)~€13.000–16.70096 GB VRAM; aber fast dein gesamtes Budget

Decision Rule: Start with 32 GB. If you hit VRAM ceiling regularly → upgrade later. Mainboard is 2-GPU-ready.

10GbE is Non-Negotiable

Parvotec generates ~1.8 TB embeddings per BO round. Without 10GbE:

Cost addition: €300–800 (cards + switch) — mandatory.

Scientific Node — Hardware Bill of Materials

Complete System (~€8–10k netto, €9.5–12k brutto)

KomponenteModellNettoBruttoLink / Quelle
CPUAMD Threadripper 9960X (24c, 5.7 GHz)€800€960Geizhals.at
MainboardASUS ProArt TRX50-CREATOR€1.100€1.32010GbE, dual PCIe x16, ECC support
RAMCrucial CT2K64G56C46S5 (2×64 GB, DDR5-5600, ECC)€1.800€2.160Alternate.de
GPUNVIDIA RTX 5090 32 GB (Founder's Edition or AIB)€4.500–5.500€5.400–6.600aktuell €5–6k, MSRP war $1.999
System SSDSamsung 990 Pro 2 TB (PCIe 4.0)€180€216OS, containers, compilers
Data SSDSK Hynix Platinum P41 4 TB (PCIe 4.0)€280€336Embeddings cache, training data, models
PSUCORSAIR AX1600i (1600 W, Titanium-rated)€400€480RTX 5090 + Threadripper = ~700 W peak; 1600 W = 2.3× headroom
CaseNZXT H7 Flow RGB oder Corsair 5000T€120€144Good airflow, quiet, full-size for later GPU expansion
Cooler (CPU)Noctua NH-U14S TR5-SP6 oder be quiet! Dark Rock Pro TR4€150€180Threadripper runs hot; high-quality cooler essential
10GbE CardMellanox ConnectX-6 (QSFP28, dual-port)€250€300Or Cisco ENIC (cheaper used)
10GbE SwitchNVIDIA Mellanox SN3700 (16-port managed)€1.500€1.800High-performance, power-efficient, redundant PSU support
Cables + MiscQSFP+ cables (2×3m), PCIe risers, thermal paste€100€120
TOTAL HARDWARE€11.180€13.416Without 10GbE switch: €9.5k netto

Optional Expansions (Für später)

Cost Analysis & Procurement Timeline

Budget Breakdown

KategorieKomponentenNettoBrutto (20% MwSt)
Core ComputeCPU + Mainboard + RAM + GPU€8.200€9.840
Storage2×NVMe (2TB + 4TB)€460€552
Power & CoolingPSU + Cooler€550€660
GehäuseCase + Cables + Misc€220€264
SYSTEM SUBTOTAL€9.430€11.316
Network (Shared)10GbE Card + Switch + Cables€1.850€2.220
TOTAL WITH 10GbE€11.280€13.536

Procurement Strategy (M0–M2)

M0 (Now)Order CPU, Mainboard, RAM, PSU (longest lead time)
M0+2wGPU pre-order (RTX 5090 supply constrained; check Geizhals daily)
M1Receive core components, assemble + BIOS setup
M1+2wGPU arrives, 10GbE NIC + Switch (can order in parallel)
M2Full system operational: Ubuntu 24.04 LTS + CUDA 12.5

Austrian Retailers

Work Package Integration — What Do You Need When?

WP1: Target Definition (M1–M3)

TaskHardware RequirementWhy Workstation?Fallback
Sequence assembly (FASTQ → HDF5)CPU (16 cores) + 32 GB RAMThreadripper 9960X handles bioinformatics pipelines faster than MacBookMacBook M5 (slower)
ESM-2 pilot embedding (100–500 seqs)GPU (RTX 5090, ~4h)5090 ~100× faster than M5 Metal; parallelize batchesMacBook M5 (~30h, offline only)
Result visualization + QCCPU onlyLocal Jupyter, metadata inspectionMacBook (preferred, but workstation OK)

WP1 Hardware Status: ✓ READY with Threadripper + RTX 5090

WP2: DMS Library (M2–M8)

TaskHardware RequirementTimelineCriticality
Full ESM-2 embedding (500k seqs)GPU 24/7 (RTX 5090, ~72h continuous)M2–M5 (first 3 weeks)🔴 CRITICAL PATH
Data staging (500k seq input)CPU + 128 GB RAM + 4 TB SSD cacheM1–M2🟡 High
Embedding validation samplingCPU + metadata QC (no GPU)M5–M8🟠 Medium
External NAS deployment10GbE network + 4 TB RAID-1 NASBefore M8 BO start🔴 CRITICAL (storage pressure)

⚠️ KEY CONSTRAINT: Workstation must have 10GbE + external NAS by M8, else 3TB local NVMe becomes bottleneck for WP4 (1.8 TB embeddings × 3 rounds).

WP3: Oracle v1.0 (M4–M10)

TaskHardwareGPU HoursWhy Workstation
4-Task NN training (4 epochs)RTX 5090 + 128 GB RAM~150 GPU-h5090 = 10–20× faster than M5 Metal; memory bandwidth critical
cVAE training (conditional VAE)RTX 5090 (lower footprint)~100 GPU-hGenerative task; M5 Metal insufficient
Hyperparameter sweep (Optuna)CPU (Threadripper 24c) + GPU~50 GPU-h (BO search)Parallelizable; Threadripper >> M5 cores
Model checkpointing + validationStorage + CPUBoth systems OK; prefer workstation for automation

WP3 Status: RTX 5090 + Threadripper ✓ SUFFICIENT for M4–M10 timeline.

WP4: BO Runden (M8–M18)

RoundWorkloadHardware BottleneckEstimated Time
Round 1 (M8–M11)Embed (72h) → NN Infer (18h) → BO GP (48h CPU)GPU embedding (RTX 5090)~8–9 calendar days (sequential)
Round 2 (M11–M14)Repeat Round 1 with 2nd candidate setGPU + 4 TB NAS (1.2 TB embeddings now)~8–9 days (may slow with NAS I/O)
Round 3 (M14–M18)Repeat Round 1 refined final setStorage pressure: ~1.8 TB cumulative~8–9 days (watch NAS bandwidth)
Total BO3 rounds sequential (no DDP parallelism)GPU sustained 24/7 × ~24 days~30 calendar days (fits M8–M18)

⚠️ CRITICAL: 10GbE + NAS Tier 2 are MANDATORY for WP4. Without them, 4TB local NVMe insufficient.

WP5: Lead Validation (M14–M24)

Wet Lab IntegrationMacBook M5 (field) + Workstation (overnight)MacBook handles field data ingest; Workstation processes batch
Transfer LearningRTX 5090 + Threadripper (re-training)~300 GPU-h spread over 10 months
Inference on New CandidatesEither system; prefer M5 for interactivityOracle inference = ~1s/seq on M5 (acceptable)

Summary: Hardware Readiness by WP

WPCore ComputeStorageNetworkStatus
WP1Threadripper + RTX 50906 TB local NVMe1GbE OK✓ Ready M2
WP2✓ (Threadripper 24c parallelism)⚠️ Need 4TB NAS by M8⚠️ Need 10GbE before BO✓ M2–M8 feasible
WP3✓ (RTX 5090 training)✓ (Tier 1 + Tier 2 warm)✓ 10GbE for Parvotec↔Mac sync✓ Ready M4
WP4✓ (RTX 5090 24/7)🔴 CRITICAL: 4TB NAS MANDATORY🔴 10GbE MANDATORY⚠️ M8 only if NAS+10GbE deployed
WP5Workstation + MacBook (hybrid)✓ Archive to S3 Glacier✓ Ready M14

M1–M24 Timeline & Workstation Usage

Monthly GPU/CPU Utilization (Workstation)

MonthWPGPU UsageCPU UsageNAS PressureNotes
M1–M2Setup + WP1 start~20% (pilot embedding)50% (sequence prep)OS setup, CUDA drivers, Python env
M2–M5WP2 main (DMS)95% (500k embedding)30% (I/O scheduling)Continuous 72h ESM-2 run; local cache OK
M4–M8WP3 overlap60% (NN training)80% (hyperparameter BO)Both systems busy; parallelizable
M8WP2 complete🔴 NAS deployment DEADLINEMust have 4TB RAID-1 before WP4 start
M8–M11WP4.1 (Round 1)99% (BO embedding)70% (GP process)~600 GB writtenIntensive; 10GbE critical for staging
M11–M14WP4.2 (Round 2)99%70%~1.2 TB totalPotential NAS bandwidth bottleneck; monitor
M14–M18WP4.3 + WP5 start99% (BO final)50% (validation)~1.8 TB finalHottest period; M5 does WP5 field work
M18–M24WP5 main20% (transfer learn)40% (analysis)Archive to S3Workstation in idle/standby most days

Critical Dependencies

Workstation Power Consumption

Idle~150 W (mostly PSU fan)
CPU-only load~300 W (Threadripper + fans)
GPU streaming (RTX 5090)~700 W peak (Threadripper 200W + GPU 450W + PSU loss)
24/7 continuous (M8–M18)~5,040 kWh/month (worst case, avg ~600W)

Recommendation: UPS 6 kVA (€3k) highly recommended for WP4 BO rounds. Sudden power loss = 72h embedding lost.