Digital & Intelligence Engineering / 10-15+ Years
Solution Architect-AI Infrastructure Cloud
Bengaluru
Job Description:
We are looking for an experienced Solution Architect – AI Infrastructure & Private Cloud with expertise in AI/ML infrastructure, HPC, Kubernetes, and hybrid cloud platforms. The role involves architecting, deploying, and optimizing enterprise AI platforms using HPE Private Cloud AI, HPE GreenLake, and NVIDIA AI Enterprise solutions.
The ideal candidate should have strong experience in AI infrastructure, container orchestration, automation, and enterprise-scale private cloud environments.
Key Responsibilities:
Leadership & Strategy
- Lead architecture and delivery of AI/ML and HPC platform solutions aligned with customer requirements.
- Ensure scalable, secure, and high-performance AI infrastructure design.
- Manage stakeholder communication, project risks, and delivery governance.
Solution Design & Delivery
- Design and optimize solutions using Kubernetes platforms such as OpenShift, Rancher, and HPC schedulers like Slurm/PBS Pro
- Integrate AI platforms with NVIDIA AI Enterprise, DevOps, and MLOps ecosystems.
- Support hybrid cloud and GPU-accelerated workloads.
Opportunity Assessment
- Lead technical discussions for RFPs, RFIs, PoCs, and customer solution workshops.
- Assess customer environments and recommend suitable HPE and partner reference architectures.
Innovation & Collaboration
- Stay updated with emerging technologies across AI/ML, HPC, Kubernetes, and cloud-native platforms.
- Collaborate with infrastructure, networking, storage, and data science teams.
- Mentor technical teams and support knowledge-sharing initiatives.
Required Skills:
HPC & AI Infrastructure
- Experience with HPC technologies, Slurm/PBS Pro, HPCM, and NVIDIA Base Command Manager.
- Knowledge of GPU technologies, InfiniBand, Mellanox, and AI infrastructure optimization.
Containerization & Orchestration
- Strong experience with Docker, Podman, Singularity, Kubernetes, OpenShift, Rancher, and K3S.
- Experience with NVIDIA GPU Operator and DCGM.
Linux & Virtualization
- Strong Linux administration skills across RHEL, Ubuntu, and SLES.
- Experience with KVM and OpenShift Virtualization.
Cloud, DevOps & MLOps
- Experience with hybrid cloud environments, CI/CD, IaC, and MLOps workflows.
- Knowledge of observability, security, and compliance frameworks.
Networking & Automation
- Strong understanding of DNS, TCP/IP, routing, load balancing, S3, NFS, and SMB/CIFS.
- Proficiency in Python and Bash scripting for infrastructure automation.
Soft Skills
- Strong communication, problem-solving, and stakeholder management skills.
- Ability to lead enterprise-scale technical projects end-to-end.
Qualifications
- Bachelor’s or Master’s degree in Computer Science, IT, or related field.
- 10-15 years of experience in AI/ML, HPC, Kubernetes, and private/hybrid cloud environments.
- Certifications such as RHCSA, RHCE, CKA, CKAD, CKS, or NVIDIA AI certifications are preferred.