Skip to main content

Eqvilent

С++ / СUDA developer

Remote (Remote)midAdded 1 month ago

About this role

Join a dynamic team as a C++/CUDA developer to build high-performance GPU and CPU solutions with extreme latency and throughput demands. You'll implement complex computational algorithms, optimize existing systems for scalability, and work on cutting-edge low-latency projects in a fully remote environment.

What you'll do

  • Implement complex computational algorithms on GPU and CPU with strict latency and throughput requirements
  • Design and optimize custom CUDA kernels for performance
  • Refactor existing solutions to improve scalability
  • Profile and analyze application performance using specialized tools
  • Debug high-performance GPU and CPU applications
  • Collaborate with international team on project-based work

What they're looking for

  • C++ (strong proficiency with data structures, algorithms, OOP)
  • CUDA programming and custom kernel design
  • GPU/CPU performance optimization
  • Profiling tools (Nsight Systems, Nsight Compute, nvprof)
  • Low-latency and real-time system development
  • Linux system internals and networking
  • Lock-free data structures
  • Memory optimization techniques

Benefits

  • Fully remote work from anywhere globally
  • Access to global offices
  • Flexible schedule
  • 40 paid days off
  • Competitive salary
  • Cutting-edge hardware and technology
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Eqvilent

Eqvilent builds high-performance, low-latency trading infrastructure and systems that process massive financial market data volumes with extreme speed requirements. The company is hiring C++/CUDA developers, systems engineers, and software engineers to develop production trading platforms, optimize computational algorithms, and maintain globally distributed infrastructure.

View all jobs at Eqvilent

Likely interview questions

  • Describe your experience designing and optimizing custom CUDA kernels. What specific challenges did you face regarding latency and throughput, and how did you address them?
  • Walk us through your approach to profiling GPU applications. Which tools have you used (Nsight Systems, Nsight Compute, nvprof) and how did you use their results to optimize performance?