Loading jobs…
Loading jobs…
Submer — European district, Brussels Capital
About Radian Arc. Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.
What impact you will have Mission: Design, build, and operate the AI storage layer powering large-scale GPU infrastructure, enabling datasets, model artifacts, checkpoints, and inference state to be delivered to compute clusters with extremely high throughput and predictable latency. You will play a key role in architecting and evolving the storage platform across edge and core deployments, supporting the full lifecycle of AI workloads including distributed inference, fine-tuning, and large-scale model training. The role spans multiple storage architectures used across the platform, including hyperconverged storage currently based on StorPool, local NVMe storage for latency-sensitive workloads and edge deployments, and disaggregated AI storage platforms such as VAST Data and Weka.
As the first dedicated storage platform role in the organization, this position combines Staff-level architectural ownership, technical direction, and cross-functional influence with hands-on execution across storage design, deployment, performance engineering, troubleshooting, platform integration, and operational improvement. A key responsibility of this role is designing and optimizing the storage architecture underlying distributed inference stacks such as NVIDIA Dynamo, llm-d, or similar inference orchestration frameworks. This includes ensuring that storage systems efficiently support inference workloads through optimized dataset access, model artifact distribution, checkpoint handling, and KV-cache persistence.
You will design scalable storage systems capable of feeding thousands of GPUs while balancing throughput, latency, resilience, and cost efficiency, and work closely with compute, networking, and platform engineering teams to ensure seamless integration with the platform orchestration layer. Because this is currently the primary storage platform role in the company, the position is intentionally hybrid: you are expected to operate at L6 / Staff in terms of long-term design, standards, cross-team influence, and platform direction, while also directly executing critical storage work that, in a larger organization, would be distributed across multiple engineers.
What You’Ll Do
Storage Architecture Design scalable AI storage architectures supporting both edge and core deployments. Define storage strategies for distributed inference, fine-tuning, and training workloads. Architect solutions across multiple storage models: Hyperconverged infrastructure such as StorPool, Local NVMe storage, Disaggregated storage systems such as VAST, Weka, and related architectures.
Define reference architectures, design principles, and reusable patterns for storage platforms so future deployments follow standards rather than one-off implementations. Evaluate trade-offs across throughput, latency, resilience, data locality, cost, and operability, and make clear recommendations to engineering and leadership. Influence the long-term storage roadmap, including architecture choices for edge, core, hyperconverged, and disaggregated environments.
AI Workload Optimization Optimize storage throughput and latency for GPU-heavy clusters. Design data locality strategies to minimize dataset movement across the network. Benchmark storage performance under real AI workloads.
Optimize I/O patterns for large dataset ingestion, checkpointing, and model artifact distribution.