Top

Senior HPC and AI Networking Performance Research and Analysis Engineer

Santa Clara, CA, USA

278 Days ago

Job Description


Intelligent machines powered by Artificial Intelligence computers that can learn, reason and interact with people are no longer science fiction. GPU Deep Learning has provided the foundation for machines to learn, perceive, reason and solve problems. Today, visual computing is a crucial tool in helping people get along with technology, and NVIDIA has extended its technology into datacenters, mobile devices and cars. There has never been a more exciting time to join our team - if this role sounds like a fit for you, we'd love to hear from you!

NVIDIA is seeking a Senior High Performance Computing (HPC) and AI Networking Performance Research and Analysis Engineer to join our Performance group. In this exciting role, you will profile and analyze AI workloads on large GPUs and CPUs scale clusters for distributed Deep Learning LLM training focused on collectives communication and networking. You will interact with many types of hardware and platforms, such as HCAs, Switches, CPUs, GPUs, and Systems. You will develop performance analysis tools and methodologies to dive deeply into the details and understand performance expectations, limitations, and bottlenecks.

What you'll be doing:

Exploring and researching AI workloads and DL models specifically tailored for large-scale deep learning LLM training on NVIDIA supercomputers and distributed systems focusing on high-performance networking and Nvidia Collective Communications Library (NCCL).

Benchmarking, Profiling, and Analyzing the performance to find bottlenecks and identify areas of improvement and optimizations, with a strong emphasis on networking aspects.

Implementing performance analysis tools.

Collaborating with many teams from hardware to software to provide performance analysis insights.

Defining performance test planning , setting performance expectations for new technologies and solutions, and working to reach the performance targets limits.

What we need to see:

B.Sc in Computer Science or Software Engineering or equivalent experience

5+ years of experience with high-performance Networking (RDMA, MPI, NCCL, Congestion Control Algorithms)

Demonstrated Performance Analysis skills and methodologies.

Experience with NVIDIA GPUs, CUDA library, deep learning frameworks like TensorFlow or PyTorch, combined with expertise in networking collective communication libraries (such as NCCL) and protocols (such as RoCE and RDMA).

Fast and self-learning capabilities with strong analytical and problem-solving skills.

Programming Languages: Python, Bash and C languages

Experience with Linux OS distros.

Great teammate with good communication and interpersonal skills

Ways to stand out from the crowd:

In-depth knowledge and experience with AI workloads and benchmarking for distributed LLM training.

Knowledge in CUDA, and NCCL libraries.

Knowledge in Congestion Control algorithms.

In-depth System knowledge and understanding (Intel / AMD / ARM CPUs, NVIDIA GPUs, HCA, Memory, PCI).

Strong Performance Analysis skills and methodologies using modern tools.

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. We have a unique legacy of innovation that's fueled by great technology?and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what?s never been done before takes vision, innovation, and the world's best talent. Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As an NVIDIAN, you'll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world!

#LI-Hybrid

The base salary range is 148,000 USD - 287,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits (https://www.nvidia.com/en-us/benefits/) .

NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Key Skills Required

PythonNetworkingLinux OSAlgorithmsComputer GraphicsAnalysisArtificial IntelligenceBashBenchmarkingCommunicationComprehensiveComputer ScienceComputingCUDADeep LearningFocusedHigh Performance ComputingInnovationIntelligenceInterpersonal SkillsLearningLinuxOrientationPerformance AnalysisProfilingPyTorchResearchScienceSoftware EngineeringSupportiveTappingTensorFlowTest PlanningTraining

Job Overview


Job Function: Other

Job Type: Full Time

Workplace Type: Not Specified

Experience Level: Mid-Senior level

Salary: Competitive & Based on Experience

Experience: 5 - 6 yrs

Contact Information


Company about us:

NVIDIA is a leading company in the world of accelerated computing, with a rich history of innovation and growth. Since its inception in 1993, the company has been at the forefront of revolutionizing the way we use technology, particularly in the fields of gaming, computer graphics, and artificial intelligence. With...

Company Name: NVIDIA

Recruiting People: HR Department

Website: https://www.nvidia.com/en-us/

Headquarter: Santa Clara, California, USA 95050

Industry: IT/Computers - Hardware & Networking

Company Size: 10000+ Employees

Location

Important Fraud Alert:
Beware of imposters. elsejob.com does not guarantee job offers or interviews in exchange for payment. Any requests for money under the guise of registration fees, refundable deposits, or similar claims are fraudulent. Please stay vigilant and report suspicious activity.

Similar Jobs

Medical Visualization R&D Manager

Anatomage, Inc. • Santa Clara, CA, USA

Experience: 5 - 7 yrs

Salary: Competitive & Based on Experience

View Job
Optical Engineering Manager

Halo Industries, Inc. • Santa Clara, CA, USA

Experience: 10 - 11 yrs

Salary: $175,000 - $190,000 / Annual Salary

View Job
Quality Manager (Engineer background)

T45 Labs • Santa Clara, CA, USA

Salary: $118,000 - $160,000 / Annual Salary

View Job
Conserje/Janitor

Impec Group • Santa Clara, CA, USA

Salary: Competitive & Based on Experience

View Job
Digital Science Content Specialist

Anatomage, Inc. • Santa Clara, CA, USA

Salary: Competitive & Based on Experience

View Job
Staff Process Engineer (Wafer Finishing)

Halo Industries, Inc. • Santa Clara, CA, USA

Experience: 5 - 6 yrs

Salary: $155,000 - $170,000 / Annual Salary

View Job
Senior NPI Engineer

Halo Industries, Inc. • Santa Clara, CA, USA

Experience: 5 - 6 yrs

Salary: $160,000 - $180,000 / Annual Salary

View Job
Nationwide Strategic Account Manager

Anatomage, Inc. • Santa Clara, CA, USA

Experience: 3 - 5 yrs

Salary: Competitive & Based on Experience

View Job
Director of New Product Introduction (NPI)

Halo Industries, Inc. • Santa Clara, CA, USA

Salary: $200,000 - $220,000 / Annual Salary

View Job
Nationwide Account Executive

Anatomage, Inc. • Santa Clara, CA, USA

Experience: 3 - 5 yrs

Salary: Competitive & Based on Experience

View Job