×
Register Here to Apply for Jobs or Post Jobs. X

Machine Learning Engineer - Machine Learning Infrastructure

Job in San Jose, Santa Clara County, California, 95199, USA
Listing for: ByteDance
Full Time position
Listed on 2026-03-01
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, Software Engineer, Data Engineer
Salary/Wage Range or Industry Benchmark: 187040 - 438000 USD Yearly USD 187040.00 438000.00 YEAR
Job Description & How to Apply Below

Machine Learning Engineer - Machine Learning Infrastructure

Location:

San Jose

Team:
Technology

Employment Type:

Regular

Job Code: JLKPV

Responsibilities
  • About the Team:
    The mission of our AML team is to push next-generation machine learning algorithms and platforms for the recommendation system, ads ranking and search ranking in our company. We also drive substantial impact on core businesses of the company.
  • Resource Efficiency Optimization in Distributed Orchestration and Scheduling:
    Develop and extend distributed orchestration frameworks within the Kubernetes/Godel ecosystem. Select appropriate frameworks based on different business scenarios, and optimize cluster utilization and load balancing strategies according to the specific characteristics of each scenario;
    Integrate and expand Auto Scaling and automatic parallelization capabilities for various models and tasks. Employ load modeling and analytic methods for different models to automatically optimize resource requests, achieving large-scale improvements in resource usage efficiency and global optimality;
    Responsible for preemption and re-scheduling mechanisms for services with different priorities, and manage automatic resource multiplexing across different clusters and resource types; handle scheduling and load adaptation across multi-datacenter, multi-region, and multi-cloud environments.
  • Building Training System Architecture for Next-Generation Ultra-Large and Ultra-Deep Recommendation Models:
    Develop a flexible, elastic and robust distributed training runtime focused on hyper-scaled embeddings and large-scale GPU training;
    Design and optimize distributed computing APIs and runtimes geared towards future recommendation and ads model paradigms (e.g., reinforcement learning, fine-tuning and/or distillation);
    Collaborate with platform teams to enhance the diagnosability and usability of distributed training systems.
  • Constructing Online Orchestration Architecture for Next-Generation Recommendation Systems:
    Build a robust and stable distributed model inference architecture for online learning scenarios involving hyper-scaled embeddings;
    Optimize the usability of online recommendation and ads model architectures and MLops workflows.
Qualifications

Required Qualifications

  • Bachelor's degree or above, majoring in Computer Science, Engineering or related fields.
  • Strong programming and coding experience with at least one modern language such as Golang, Python
  • Experience contributing to the large scale distributed systems, multi-tenant systems (architecture, reliability and scaling)
  • Strong analytical abilities and problem solving
  • Good communication, self-motivation, engineering practice, documentation, etc.
  • At least 3 years of relevant experience.

Preferred Qualifications

  • Familiar with large-scale distributed scheduling systems like Kubernetes, Yarn, Flink and/or Spark;
  • Familiar with opensourced orchestration frameworks like VeRL, vLLM, Ray or TFX, etc.;
Job Information

The base salary range for this position in the selected city is $187040 - $438000 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

Legal and Diversity

For Los Angeles County (unincorporated) Candidates:
Qualified applicants with arrest or conviction records will…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)

Job Posting Language
Employment Category
Education (minimum level)
Filters
Education Level
Experience Level (years)
Posted in last:
Salary