Employer profile
Inferact
5 open roles indexed with location, benefit, and apply-link signals where available.
Open roles
Showing the most recent indexed roles for this employer.
-
Member of Technical Staff, Exceptional Generalist (Remote)
Remote, US
remote Salary not disclosedMember of Technical Staff, Exceptional Generalist (Remote) Remote, US Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build. About the Role This is a globally remote opportunity. We're seeking exceptional generalist engineers who can work across the entire vLLM stack: from low-level GPU kernels to high-level distributed systems. This role is designed for self-directed, autonomous individuals who can identify the highest-leverage problems and solve them end-to-end without constant guidance. You'll work asynchronously with our San Francisco headquarters while maintaining full ownership of critical infrastructure. You might be optimizing CUDA kernels one week, designing distributed orchestration systems the next, and implementing new model architectures the week after. The work you do will directly impact how the world runs AI inference. Potential focus areas include: - Inference Runtime: Push the boundaries of LLM and diffusion model serving. Work at the core of vLLM to optimize how models execute across diverse hardware and architectures. - Kernel Engineering: Write the low-level kernels and optimizations that make vLLM the fastest inference engine in the world, running on hundreds of accelerator types. - Performance & Scale: Build the distributed systems that power inference at global scale-design foundational layers enabling vLLM to serve models across thousands of accelerators with minimal latency. - Cloud Orchestration: Build the operational backbone for cluster
-
Member of Technical Staff, Cloud Orchestration
San Francisco, California, United States
unspecified $200K-$400KMember of Technical Staff, Cloud Orchestration San Francisco, California, United States Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build. About the Role We're looking for an cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at massive scale. You'll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction. You'll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works. Skills and Qualifications Minimum qualifications: - Bachelor's degree or equivalent experience in computer science, engineering, or similar. - Strong experience with Kubernetes and container orchestration at scale. - Experience designing and implementing custom Kubernetes operators. - Proficiency in Python/Rust/Go and infrastructure-as-code tools (Terraform, Helm, etc). - Experience managing GPU clusters and debugging hardware issues. - Ability to work across cloud platforms (AWS, GCP, Azure) and on-premise infrastructure. Preferred qualifications: - Experience with ML-specific orchestration tools (Ray, Slurm). - Knowledge of GPU scheduling, multi-tenancy, and resource optimization. - Familiarity with vLLM deployment patterns and configuration. - Track record of improving operational reliability for ML systems. Bonus points if you have: - Experience deploying inference systems on large-scale GPU (1,000+) clusters. Logistics - Location: This role is based in San Francisco, California.
-
Member of Technical Staff, Kernel Engineering
San Francisco, California, United States
unspecified $200K-$400KMember of Technical Staff, Kernel Engineering San Francisco, California, United States Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build. About the Role We're looking for a performance engineer to squeeze every FLOP out of modern accelerators. You'll write the kernels and low-level optimizations that make vLLM the fastest inference engine in the world. Your code will run on hundreds of accelerator types, from NVIDIA GPUs to emerging silicon. When hardware vendors develop new chips, they integrate with vLLM. You'll work directly with these teams to ensure we're extracting maximum performance from every generation of hardware. Skills and Qualifications Minimum qualifications: - Bachelor's degree or equivalent experience in computer science, engineering, or similar. - Deep experience writing CUDA kernels or equivalent (CuTeDSL, Triton, TileLang, Pallas). - Strong understanding of GPU architecture: memory hierarchy, warp scheduling, tiling, tensor cores. - Proficiency in C++ and Python with demonstrated ability to write high-performance code. - Experience with profiling tools (Nsight, rocprof) and performance optimization methodologies. - Obsession with benchmarks and squeezing every percentage point of speedup. Preferred qualifications: - Experience with ML-specific kernel optimization (FlashAttention, fused kernels). - Knowledge of quantization techniques (INT8, FP8, mixed-precision). - Familiarity with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel). - Experience with compiler technologies (LLVM, MLIR, XLA). Bonus points if
-
Member of Technical Staff, Performance and Scale
San Francisco, California, United States
unspecified $200K-$400KMember of Technical Staff, Performance and Scale San Francisco, California, United States Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build. About the Role We're looking for an infrastructure engineer to build the distributed systems that power inference at global scale. You'll design and implement the foundational layers that enable vLLM to serve models across thousands of accelerators with minimal latency and maximum reliability. Tomorrow, deploying a frontier model at scale should be as straightforward as spinning up a serverless database. The complexity doesn't disappear as it gets absorbed into the infrastructure you're building. Skills and Qualifications Minimum qualifications: - Bachelor's degree or equivalent experience in computer science, engineering, or similar. - Strong systems programming skills in Rust, Go, or C++. - Experience designing and building high-performance distributed systems at scale. - Understanding of network protocols and high-performance I/O. - Ability to debug complex distributed systems issues. Preferred qualifications: - Experience with ML serving infrastructure and disaggregated inference architecture. - Familiarity with GPU programming models and memory hierarchies. - Knowledge of GPU interconnects (NVLink, InfiniBand, RoCE) and their performance characteristics. - Track record of improving system reliability and performance at scale. Bonus points if you have: - Prior experience in supporting large‑scale model training or inference environments. Logistics - Location: This role is
-
Member of Technical Staff, Inference
San Francisco, California, United States
unspecified $200K-$400KMember of Technical Staff, Inference San Francisco, California, United States Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build. About the Role We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference. Skills and Qualifications Minimum qualifications: - Bachelor's degree or equivalent experience in computer science, engineering, or similar. - Deep understanding of transformer architectures and their variants. - Strong programming skills in Python with experience in PyTorch internals. - Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI). - Ability to read and implement model architectures and inference techniques from research papers. - Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases. Preferred qualifications: - Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving. - Familiarity with RL frameworks and algorithms for LLMs. - Experience with multimodal inference (audio/image/video/text). - Contributions to open-source ML or system infrastructure projects. Bonus points if you have: - Implemented core features in vLLM