SuperQ Quantum Computing logo

AI Engineer (GPU & LLM Optimisation) – Canada / UAE / Remote

SuperQ Quantum Computing

Actively hiring🇨🇦 Remote (Canada)RemoteFull-timeC$90k – C$120k /yr2–4 yrs expData Scientist
Posted 2d agoBe an early applicant

About the Role SuperQ is looking for an AI Engineer specializing in GPU and LLM optimization to maximize the performance, scale, and efficiency of the models powering our multi-agent architecture. You will be the driving force behind accelerating inference, reducing latency, and optimizing hardware utilization as we scale our hybrid compute platform across industries like healthcare, manufacturing, and finance.

Responsibilities

  • Accelerate Inference: Optimize LLM inference and serving pipelines using high-performance frameworks (e.g., vLLM, TensorRT-LLM, TGI, Triton Inference Server).
  • Model Compression: Implement state-of-the-art model compression techniques, including quantization (e.g., GPTQ, AWQ, FP8/INT8), pruning, and knowledge distillation to reduce memory footprints.
  • Hardware Profiling: Profile GPU performance to identify and eliminate bottlenecks, optimizing memory management techniques such as KV caching and PagedAttention.
  • Custom Compute: Develop and optimize custom CUDA or OpenAI Triton kernels for specialized, computationally heavy operations where standard libraries fall short.
  • Cross-functional Deployment: Collaborate with backend and platform engineering teams to deploy highly scalable, low-latency AI endpoints within our Kubernetes-based infrastructure.
  • System Monitoring: Ensure robust monitoring of GPU health, utilization metrics, latency, and throughput in production environments.

Requirements

  • Experience: 2-4 years of experience in ML engineering, AI infrastructure, or high-performance computing (HPC) with a strong focus on Large Language Models.
  • Core Languages: Strong programming skills in Python; proficiency in C++ and/or CUDA is highly preferred.
  • Frameworks: Hands-on experience with LLM serving architectures, distributed computing, and optimization libraries (e.g., DeepSpeed, Hugging Face Accelerate, Ray).
  • Hardware Knowledge: Deep understanding of GPU architectures (NVIDIA), memory hierarchies, and parallel computing paradigms (e.g., NCCL).
  • Infrastructure: Solid understanding of containerization and orchestration (Docker, Kubernetes) tailored for GPU-accelerated workloads.
  • Bonus: Experience with hardware-aware neural architecture search (NAS), ML Ops, or working at the intersection of classical AI hardware and quantum compute interfaces.

Benefits

  • Competitive salary plus bonus based on module delivery and platform impact.
  • Stock options / equity - participate in the growth of a foundational tech platform.
  • Global remote-friendly roles: Canada, UAE or anywhere remote; with occasional team meetups and hack-weeks.
  • Learning & development stipend: conferences, AI/LLM workshops, certifications (e.g., in ML Ops, prompt engineering).
  • "Innovation time" built into schedule: work on passion projects, internal hackathons, share learnings with the team.
  • Cross-domain exposure: work across industries, build modules for different verticals - variety and career growth built-in.

Responsibilities

  • Accelerate Inference: Optimize LLM inference and serving pipelines using high-performance frameworks (e.g., vLLM, TensorRT-LLM, TGI, Triton Inference Server).
  • Model Compression: Implement state-of-the-art model compression techniques, including quantization (e.g., GPTQ, AWQ, FP8/INT8), pruning, and knowledge distillation to reduce memory footprints.
  • Hardware Profiling: Profile GPU performance to identify and eliminate bottlenecks, optimizing memory management techniques such as KV caching and PagedAttention.
  • Custom Compute: Develop and optimize custom CUDA or OpenAI Triton kernels for specialized, computationally heavy operations where standard libraries fall short.
  • Cross-functional Deployment: Collaborate with backend and platform engineering teams to deploy highly scalable, low-latency AI endpoints within our Kubernetes-based infrastructure.
  • System Monitoring: Ensure robust monitoring of GPU health, utilization metrics, latency, and throughput in production environments.

Requirements

  • Experience: 2-4 years of experience in ML engineering, AI infrastructure, or high-performance computing (HPC) with a strong focus on Large Language Models.
  • Core Languages: Strong programming skills in Python; proficiency in C++ and/or CUDA is highly preferred.
  • Frameworks: Hands-on experience with LLM serving architectures, distributed computing, and optimization libraries (e.g., DeepSpeed, Hugging Face Accelerate, Ray).
  • Hardware Knowledge: Deep understanding of GPU architectures (NVIDIA), memory hierarchies, and parallel computing paradigms (e.g., NCCL).
  • Infrastructure: Solid understanding of containerization and orchestration (Docker, Kubernetes) tailored for GPU-accelerated workloads.
  • Bonus: Experience with hardware-aware neural architecture search (NAS), ML Ops, or working at the intersection of classical AI hardware and quantum compute interfaces.

Skills

Benefits

  • Competitive salary plus bonus based on module delivery and platform impact.
  • Stock options / equity - participate in the growth of a foundational tech platform.
  • Global remote-friendly roles: Canada, UAE or anywhere remote; with occasional team meetups and hack-weeks.
  • Learning & development stipend: conferences, AI/LLM workshops, certifications (e.g., in ML Ops, prompt engineering).
  • "Innovation time" built into schedule: work on passion projects, internal hackathons, share learnings with the team.
  • Cross-domain exposure: work across industries, build modules for different verticals - variety and career growth built-in.