JDXpert Jobs
     
HRTMS Job Description Management

Senior HPC Storage Engineer

CMPTL & DATA SCI RSH SPEC 4 RP (005240)

UCPath Position ID: TBD_1000126

 

 

 

Position Description History/Status

Approved Date:

7/22/2026 9:41:21 PM

Date Last Edited:

7/22/2026 9:41:18 PM

Last Action Effective Date:

 

Organization Details

Business Unit (Location):

LACMP

Organization Code:

6600O

Organization:

OFFICE OF ADVANCED RESEARCH COMPUTING  

Division Code:

6610D

Division:

OFFICE OF ADVANCED RESEARCH COMPUTING  

Department:

220000 - RESEARCH TECHNOLOGY GROUP

Position Details

UCPath Position Number:

TBD_1000126

Position Description ID

262579

UC Payroll Title:

CMPTL & DATA SCI RSH SPEC 4 RP (005240)

Personnel Program

Management and Senior Professional (MSP)

Salary Grade:

STEPS

Job Code FLSA:

Exempt

Union Code (Collective Bargaining Unit):

RP: Research and Public Service PR

Employee Relations Code:

E: All Others - Not Confidential

Employee Class (Appt Type):

2 - Staff: Career

Full-Time Equivalent (FTE)

1

SUPERVISION

UCPath Reports to Position Number:

41148377

Reports to Payroll Title:

IT ARCHITECT 5 TX

UCPath Department Head Position Number:

41138579

Department Head Payroll Title:

IT ARCHITECT MGR 2


Level of Supervision Received

DIRECTION - Indicates that the incumbent establishes procedures for attaining specific goals and objectives in a broad area of work. Only the final results of work done are typically reviewed. Incumbent typically develops procedures within the limits of established policy guidelines.


POSITION SUMMARY

The Senior HPC Storage Engineer serves as a senior technical expert responsible for the architecture, design, deployment, and ongoing operation of large-scale data storage platforms that serve both campus-wide research storage needs and federated storage services spanning multiple geographically distributed high-performance computing (HPC) sites and cloud platforms. A major focus of this position is supporting iDLab, a federated research computing and data infrastructure effort spanning multiple NSF-supported HPC centers and cloud platforms. Reporting to the Chief Research Data Architect, the incumbent leads the full technical lifecycle of these storage platforms, from technology evaluation and architecture design through procurement support, deployment, production operations, performance optimization, capacity planning, and continuous evolution.

 

The Senior HPC Storage Engineer leads rigorous comparative analysis of current and emerging scale-out and federated storage technologies, including parallel filesystems, object storage systems, and data orchestration platforms; characterizes candidate solutions across performance, consistency, latency, scalability, operational complexity, cost, and failure-mode behavior; and produces empirically grounded recommendations to guide multi-year deployment decisions. Following architecture decisions, the incumbent leads implementation and carries primary operational engineering responsibility for the resulting platforms, in coordination with OARC systems staff and partner-site technical teams.

 

Operational responsibilities include deployment, configuration, monitoring, performance tuning, capacity planning, lifecycle management, security hardening, backup and recovery, and incident response for petabyte-class storage systems supporting the campus research community and multi-site research collaborations. The incumbent will design and execute benchmarking studies against representative scientific workloads, including interactive analysis, batch processing, AI/ML training pipelines, and visualization workflows, and will continuously evolve the platforms based on operational experience, community needs, and emerging technologies.

 

The Senior HPC Storage Engineer will additionally design and operate campus-wide research data storage services, including tiered storage architectures, shared and dedicated storage allocations, data movement services, and integration with research workflows. The incumbent assesses integration requirements with local, campus-wide, and federated identity systems and compute job orchestration platforms, evaluates caching strategies and consistency models appropriate for cross-site access, and engages with system administrators, platform developers, campus IT partners, and research stakeholders to deploy and ensure that storage services meet evolving research and educational needs.

 

The position requires deep technical expertise in scale-out and federated storage architectures, hands-on experience operating large-scale parallel, distributed, and object storage systems, and the ability to communicate complex technical trade-offs to both engineering and executive audiences. Excellent written and oral communication skills are required. The incumbent must be comfortable working at multiple levels of the stack, from architectural strategy through hands-on production troubleshooting.

 

Type of Supervision Received/Exercised:

Under the direction of Chief Research Data Architect, work is assigned at the project and service levels and reviewed periodically for quality, innovation, technical rigor, operational reliability, and the timely achievement of project milestones and service commitments. The incumbent is expected to exercise independent judgment and technical creativity to identify problems, recommend and implement solutions, and meet organizational goals. The Senior HPC Storage Engineer may lead technical working groups, mentor junior staff, and serve in an advisory capacity on storage matters across the organization.

 

 


Department Summary

Advanced Research Computing melds expert staff and technical infrastructure to amplify and accelerate the impact of UCLA research in the age of networked data and computation. OARC’s expertise and resources are available to all UCLA researchers engaged in digital research and scholarship. We work with faculty, student, and postdoctoral researchers; instructors; and staff and administrators.

 

OARC is a relationship-building organization. We enable digital scholarship through collaborations, partnerships, and networked communities to advance cutting-edge research capabilities at UCLA and beyond. OARC supports and enhances the university mission of education, research, and service through the development and execution of innovative and sustainable technology practices, programs, services, infrastructure, policies, and partnerships.

 


Key Responsibilities and Essential Functions

Function

Responsibilities

% Time

Storage Architecture Research, Evaluation, and Design

A-1. Survey and evaluate current and emerging storage technologies suitable for campus-scale and multi-site federated namespaces, including parallel filesystems (such as Lustre, VAST Data, pNFS, GPFS/Spectrum Scale, BeeGFS, Ceph), data orchestration platforms (such as Hammerspace), and object storage systems (such as MinIO and S3-compatible solutions at scale). (E)

A-2. Maintain current knowledge of newer and emerging technologies including DAOS, WekaFS, and other scale-out storage solutions relevant to research computing. (E)

10%

Storage Architecture Research, Evaluation, and Design

A-3. Characterize candidate solutions across performance, consistency, latency, scalability, operational complexity, cost, and failure-mode behavior; produce comparative analysis reports with quantitative benchmarks and recommended architecture with explicit tradeoff documentation. (E)

A-4. Design the target architecture for campus-wide research storage and federated multi-site storage services, including topology, data placement strategies, tiered storage layers, caching, and integration points with compute, network, and identity systems. (E)

 

10%

Storage Architecture Research, Evaluation, and Design

A-5. Develop technical specifications and phased implementation roadmaps; recommend procurement strategies and vendor engagement approaches. (E)

A-6. Assess integration requirements with local, campus-wide and federated identity systems (such as Active Directory, LDAP, CILogon, OIDC, SAML, Globus Auth). (E)

 

5%

Storage Platform Operations and Lifecycle Management

B-1. Deploy, configure, and operate large-scale parallel, distributed, and object storage systems supporting campus research and multi-site federated services. (E)

B-2. Monitor storage system performance, capacity, and health; implement and maintain observability tooling for real-time tracking and trend analysis; respond to alerts and incidents. (E)

B-3. Perform complex filesystem upgrades, kernel patches, firmware updates, and hardware refreshes with minimal disruption to running services. (E)

B-4. Plan and execute storage capacity expansion, lifecycle replacement, and end-of-life decommissioning aligned with research demand and budget cycles. (E)

 

15%

Storage Platform Operations and Lifecycle Management

B-5. Implement and maintain backup, replication, and disaster recovery strategies appropriate to the criticality of each data tier. (E)

B-6. Collaborate with security teams to implement best practices for storage deployment, identity management, access control, and software updates; ensure compliance with applicable policies. (E)

B-7. Participate in on-call rotation as required for emergency hardware and software support. (E)

B-8. Develop and maintain operational documentation, runbooks, and standard operating procedures. (E)

B-9. Diagnose and tune storage performance across client nodes, metadata services, storage servers, object services, network fabrics, RDMA-capable interconnects, routing, and protocol layers. (E)

 

15%

Campus-Wide Research Storage Service Development and Operation

C-1. Design, develop, and operate campus-wide research data storage services, including tiered storage offerings, shared and dedicated allocations, data movement services, and integration with research workflows. (E)

C-2. Engage with campus research communities to understand evolving storage needs and translate them into service offerings and capacity plans. (E)

C-3. Provide technical input into service definitions, allocation policies, and recharge or chargeback structures as applicable, in coordination with administrative leadership. (E)

C-4. Coordinate with campus IT partners and infrastructure groups to integrate research storage services with broader campus networking, identity, and data center resources. (E)

 

10%

Campus-Wide Research Storage Service Development and Operation

C-5. Evaluate and integrate new service capabilities such as collaborative sharing, publication-ready archival, and integration with cloud-based workflows. (E)

C-6. Provide escalation support for researchers facing complex I/O patterns, large-scale data movement challenges, or unusual storage access requirements. (E)

C-7. Design, operate, and troubleshoot large-scale data movement workflows using tools such as Globus, rsync, rclone, S3-compatible transfer tools, and site-to-site transfer services. (E)

10%

Benchmarking, Performance Engineering, and Empirical Validation

D-1. Design and execute benchmark studies against representative scientific workloads, including interactive analysis, batch processing, AI/ML training pipelines, and visualization workflows. (E)

D-2. Tune I/O performance for large-scale HPC and AI workloads at the filesystem, network fabric, and client levels. (E)

D-3. Analyze benchmark and production telemetry to identify performance bottlenecks, scalability constraints, and architectural tradeoffs; implement targeted optimizations. (E)

D-4. Validate vendor claims and technical specifications through empirical testing in representative environments. (E)

D-5. Document benchmark methodology, results, and interpretations in technical reports suitable for engineering and executive audiences. (E)

 

10%

Stakeholder Engagement, Technical Communication, and User Support

E-1. Present technical findings, service plans, and operational status to organizational leadership, including non-technical stakeholders, with appropriate framing for each audience. (E)

E-2. Engage with system administrators, platform developers, campus IT partners, and research stakeholders to validate technical assumptions and operational feasibility. (E)

E-3. Provide expert consultation to faculty, students, and research staff on storage architecture decisions for their projects. (E)

 

 

5%

Stakeholder Engagement, Technical Communication, and User Support

E-4. Develop user-facing documentation, training materials, and self-service tooling to reduce support burden and improve user experience. (E)

E-5. Lead technical working groups and meetings for the purpose of resolving divergent points of view on complex technical issues. (E)

E-6. Represent the organization at relevant technical workshops, conferences, vendor briefings, and peer-institution collaborations. (M)

5%

Maintain and Develop Knowledge and Skills

F-1. Maintain current knowledge of trends, best practices, and emerging technologies in research storage, federated data systems, and campus-scale data services. (E)

F-2. Maintain current knowledge of relevant industry benchmarking practices and reference workloads. (E)

F-3. Maintain current knowledge of standards bodies, open-source communities, and vendor ecosystems relevant to scientific data infrastructure. (E)

F-4. Attend relevant workshops, conferences, and technical briefings as appropriate to the role. (E)

F-5. Continue professional development through self-directed study and engagement with peer institutions and research computing communities. (E)

5%


Other Requirements - Applies to all Positions

•

Performs other duties as assigned.

•

Complies with all policies and standards.

•

Complies with the University of California, Los Angeles (UCLA) Principles of Community.

•

This position description is not intended to be a complete list of all responsibilities, duties or skills required for the job and is subject to review and change at any time, with or without notice, in accordance with the needs of the organization.


QUALIFICATIONS


Educational Requirements

Education Level

Education Details

Required/
Preferred

And/Or

Bachelor's Degree

Bachelor's degree in Computer Science, Computational Science, Data Science, Engineering, or a related field; or equivalent combination of education and experience.

Required

 

Master's Degree

Master's in Computer Science, Computational Science, Data Science, Engineering, or a related field; or equivalent advanced professional experience.

Preferred

Or

PhD

PhD in Computer Science, Computational Science, Data Science, Engineering, or a related field; or equivalent advanced professional experience.

Preferred

 


Experience Requirements

Experience

Experience Details

Required/
Preferred

And/Or

7 years or more

Experience in research, enterprise, or hyperscale storage environments with responsibility for large-scale production storage services.

Required

 

Large scale storage

Experience with petabyte-scale storage, federated storage, multi-site research infrastructure, or HPC/cloud-integrated storage services.

Preferred

 


Knowledge, Skills and Abilities

KSAs

Required/
Preferred

1. Advanced knowledge of high-performance computing, data science, and cyberinfrastructure environments supporting large-scale research workloads

Required

2. Demonstrated knowledge of scale-out, parallel, distributed, object, and federated storage architectures, including performance characteristics, consistency models, caching strategies, and operational tradeoffs.

Required

3. Hands-on experience architecting, deploying, operating, or substantially supporting one or more large-scale storage platforms such as Lustre, VAST Data, GPFS/Spectrum Scale, Ceph, BeeGFS, WekaFS, or MinIO. Experience with multiple platforms and petabyte-scale deployments is preferred.

Required

4. Demonstrated experience operating large-scale production storage systems, including monitoring, capacity planning, lifecycle management, upgrades, performance tuning, incident response, backup, replication, and disaster recovery.

Required

5. Advanced Linux systems administration skills, including kernel-level troubleshooting, performance profiling, storage hardware diagnostics, and tuning across InfiniBand, RoCE, and Ethernet fabrics.

Required

6. Demonstrated experience diagnosing storage performance problems across Linux clients, metadata services, storage servers, object services, network fabrics, protocol layers, and application I/O patterns.

Required

7. Proficiency in scripting and automation using Bash, Python, or similar languages; familiarity with configuration management tools such as Ansible and version control such as Git.

Required

8. Experience integrating storage systems with local, campus-wide, and federated identity providers such as Active Directory, LDAP, CILogon, OIDC, SAML, or Globus Auth.

Required

9. Demonstrated ability to design and execute storage benchmarks, validate vendor claims, characterize representative scientific workloads, and document results for engineering and executive audiences.

Required

10. Demonstrated ability to design, operate, or support shared research storage services, including allocation models, user-facing service delivery, large-scale data movement, and escalation support for complex research workflows.

Preferred

11. Demonstrated ability to communicate complex technical information clearly to technical staff, researchers, leadership, vendors, and external research and education audiences.

Required

12. Demonstrated ability to work independently and collaboratively, manage competing priorities, lead technical working groups, and sustain production service quality while delivering multi-month projects.

Required


SPECIAL REQUIREMENTS AND/OR CONDITIONS OF EMPLOYMENT


Reporting and Background Check Requirements

Background Check: Continued employment is contingent upon the completion of a satisfactory background investigation.

Live Scan Background Check: A Live Scan background check must be completed prior to the start of employment.


Other Special Conditions of Employment

List the other special conditions of employment for this position.

Description

Required/
Preferred

This position will consider both hybrid and fully remote candidates. On-site presence may be required as needed based on operational, maintenance, or project requirements. Participation in an on-call rotation for emergency hardware and software support is required. Occasional evenings and weekends may be required to support maintenance windows or respond to incidents.

Required


LOCATION AND PHYSICAL, ENVIRONMENTAL, MENTAL (PEM) REQUIREMENTS

Environment and Work Location Information

Environment Type:

Non-Clinical Setting

Location Setting:

Campus

Location:

Math Sciences Building


Physical Requirements

The physical requirements described here are representative of those that must be met by an employee to successfully perform the essential functions of this position.

Physical Requirements

Never

0 Hours

Occasional

Up to 3 Hours

Frequent

3 to 6 Hours

Continuous

6 to 8+ Hours

Is Essential

Standing/Walking

 

 

X

 

 

Sitting

 

 

X

 

 

Bending/Stooping

 

X

 

 

 

Squatting/Kneeling

 

X

 

 

 

Climbing

X

 

 

 

 

Lifting/Carrying/Push/Pull 0-25 lbs

 

X

 

 

 

Lifting/Carrying/Push/Pull 26-50 lbs

X

 

 

 

 

Lifting/Carrying/Push/Pull over 50 lbs

X

 

 

 

 

Physical requirements other

X

 

 

 

 


Environmental Requirements

The environmental requirements described here are representative of those that must be met by an employee to successfully perform the essential functions of this position.

Exposures

Never

0 Hours

Occasional

Up to 3 Hours

Frequent

3 to 6 Hours

Continuous

6 to 8+ Hours

Is Essential

Chemicals, dust, gases, or fumes

X

 

 

 

 

Loud noise levels

X

 

 

 

 

Marked changes in humidity or temperature

X

 

 

 

 

Microwave/Radiation

X

 

 

 

 

Operating motor vehicles and/or equipment

X

 

 

 

 

Exposures other

X

 

 

 

 


Mental Requirements

The mental requirements described here are representative of those that must be met by an employee to successfully perform the essential functions of this position.

Exposures

Never

0 Hours

Occasional

Up to 3 Hours

Frequent

3 to 6 Hours

Continuous

6 to 8+ Hours

Is Essential

Sustained attention and concentration

 

 

X

 

X

Complex problem solving/reasoning

 

 

X

 

X

Ability to organize & prioritize

 

 

X

 

X

Communication skills

 

 

X

 

X

Numerical skills

 

X

 

 

X

Mental demands other

X

 

 

 

 


Blood/Fluid Exposure Risk

The exposure described here is what can be expected of an employee in performing the essential functions of this position.

X

Classification 3:  Position in which exposure to blood, body fluids or tissues is not part of the position description. The normal routine task involves no exposure to blood, body fluids or tissues and the employee can decline to perform tasks which involve a perceived risk without retribution.