Find your next job Search jobs now Find the right job type for you Create a job alert Explore how we help job seekers Hire talent Contract talent Permanent talent Learn how we work with you Executive search Finance and Accounting Technology Marketing and Creative Legal Administrative and Customer Support Explore consulting solutions Technology Risk, Audit and Compliance Finance and Accounting Digital, Marketing and Customer Experience Legal Operations Human Resources Discover Insights 2026 Salary Guide Demand for Skilled Talent Report Job Market Outlook Press Room Tech insights Labor market overview AI in recruiting Navigating the AI era Staffing for small businesses Cost of a bad hire Browse jobs Find your next hire Our locations

Add your latest resume to match with open positions.

8 results for Site Reliability Engineer in San Francisco, CA

Senior/ Staff Site Reliability Engineer (SRE)
  • San Francisco, CA
  • onsite
  • Permanent / Full Time
  • 180000 - 210000 USD / Yearly
  • <p>We are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco, California. This role will focus on improving service health, refining observability, and partnering with engineering teams to build systems that perform consistently under real-world demand. The ideal candidate brings deep production experience, a strong automation mindset, and a practical approach to incident response and continuous improvement.</p><p><br></p><p>Responsibilities:</p><p>• Establish measurable reliability standards for critical services by creating and maintaining service indicators, objectives, and error budget practices.</p><p>• Take ownership of production stability by monitoring uptime, latency, and availability, and driving improvements that reduce operational risk.</p><p>• Lead live incident response efforts, coordinate troubleshooting during outages, and ensure issues are resolved efficiently and thoroughly.</p><p>• Run blameless post-incident reviews, document findings clearly, and track corrective actions through completion.</p><p>• Design and enhance observability across logs, metrics, and distributed tracing using tools such as Datadog, CloudWatch, Grafana, OpenTelemetry, and Sentry.</p><p>• Improve alert quality and dashboard design so engineering teams can quickly identify meaningful system issues without unnecessary noise.</p><p>• Evaluate system behavior under load, uncover performance constraints, and recommend changes that improve scalability and resource efficiency.</p><p>• Build automation and internal tooling that streamline operational work, strengthen deployment safety, and support incident management, debugging, and capacity planning.</p><p>• Contribute to infrastructure and delivery workflows across AWS, Terraform, Ansible, Linux, and GitHub Actions with a focus on dependable releases and resilient systems.</p><p>• Partner with security and compliance stakeholders to support operational standards, audit readiness, and the integration of monitoring into broader engineering practices</p>
  • 2026-08-11T00:00:00Z
Systems Engineer
  • Martinez, CA
  • onsite
  • Temporary / Contract
  • 60 - 70 USD / Hourly
  • <p>We are looking for a Sr. Systems Engineer to support and enhance a Microsoft-focused cloud environment in Martinez, California. This Long-term Systems Engineer Contract position will play a key role in strengthening Microsoft 365 services, device management, identity administration, and cloud-based automation across the organization. The ideal Systems Engineer candidate brings strong technical depth across Azure, Intune, Entra ID, and the Power Platform, along with the ability to resolve complex issues and guide broader IT efforts. This Systems Engineer role is an onsite role out of Martinez, Ca.</p><p><br></p><p>Responsibilities:</p><p>• Oversee the performance, security, and day-to-day administration of Microsoft 365 applications, including Exchange Online, SharePoint Online, Teams, and OneDrive.</p><p>• Create and refine Intune configurations for endpoint management, compliance enforcement, and software deployment across supported devices.</p><p>• Administer identity and access controls within Entra ID, including authentication methods, privileged access workflows, conditional access, and hybrid identity support.</p><p>• Support the design, maintenance, and troubleshooting of Azure-based services that underpin the organization&#39;s cloud infrastructure.</p><p>• Develop business-focused solutions using Power Apps, Power Automate, and Power BI to improve reporting and streamline operational processes.</p><p>• Investigate high-severity incidents, perform root-cause analysis, and implement durable fixes across Microsoft 365 and Azure platforms.</p><p>• Build and maintain automation using PowerShell and Microsoft Graph integrations to reduce manual administrative effort and improve consistency.</p><p>• Produce and update technical documentation covering architecture, configurations, operational procedures, and support standards.</p><p>• Assess emerging Microsoft capabilities and licensing updates, then provide recommendations on adoption, governance, and practical use.</p><p>• Provide technical mentorship to less experienced team members and collaborate with security, compliance, and IT partners on audits, controls, and risk reduction.</p>
  • 2026-08-21T00:00:00Z
RF Test and Integration Engineer
  • San Francisco, CA
  • onsite
  • Permanent / Full Time
  • 140000 - 160000 USD / Yearly
  • We are looking for an RF Test and Integration Engineer to lead verification efforts that confirm wireless hardware performs reliably in the lab, during production validation, and at deployed sites. This position combines hands-on RF measurement, automation development, and real-world integration work to support each stage of the hardware lifecycle. Based in San Francisco, California, the role partners closely with firmware, driver, and hardware teams to strengthen product quality and deployment readiness.<br><br>Responsibilities:<br>• Conduct detailed RF performance evaluations across operating conditions such as temperature, voltage, channel width, and output power to verify hardware behavior and readiness.<br>• Perform end-to-end integration testing across the full radio path, assessing components from the chipset and front-end circuitry through to antenna performance in bench and over-the-air environments.<br>• Create and maintain Python-based automation frameworks that control RF lab equipment and support repeatable regression testing for firmware and driver updates.<br>• Investigate performance degradation by tracing issues across software, firmware, drivers, and RF hardware to identify root causes and drive resolution.<br>• Execute site surveys, support RF design planning, confirm link-budget assumptions, and oversee deployment validation at field locations.<br>• Gather and analyze operational data from installed systems to improve propagation models and refine coverage planning accuracy.<br>• Lead in-house pre-compliance testing activities, coordinate with external certification partners, and troubleshoot failures that affect regulatory approval.<br>• Support qualification processes that establish release confidence for new hardware revisions before production and field use.
  • 2026-08-11T00:00:00Z
DevOps Engineer
  • San Francisco, CA
  • onsite
  • Permanent / Full Time
  • 180000 - 215000 USD / Yearly
  • <p>We are looking for a Senior or Staff DevOps Engineer to strengthen our engineering organization in San Francisco, California. In this role, you will improve the foundation that supports application delivery, platform reliability, and developer productivity across cloud-based systems. You will work closely with engineering partners to create scalable infrastructure, streamline release processes, and build dependable operational practices that support continued growth.</p><p><br></p><p>Responsibilities:</p><p>• Architect and maintain cloud infrastructure using infrastructure-as-code practices that promote consistency, scalability, and operational simplicity.</p><p>• Create and refine automated delivery pipelines to support dependable, efficient releases across multiple services and environments.</p><p>• Establish repeatable deployment approaches for serverless applications, container-based platforms, and orchestrated workflows.</p><p>• Build internal automation and developer tools that reduce manual effort and help backend and machine learning teams work more efficiently.</p><p>• Enhance development, testing, and staging environments to improve workflow speed and reduce friction during software delivery.</p><p>• Define strong monitoring and observability practices across logs, metrics, and tracing to improve system visibility and troubleshooting.</p><p>• Investigate reliability, latency, and performance issues in distributed environments and partner with engineers to implement lasting improvements.</p><p>• Strengthen operational readiness by supporting incident response, contributing to post-incident reviews, and turning findings into better tooling and processes.</p><p>• Integrate security and compliance controls into infrastructure and deployment workflows, including access management, patching, and vulnerability remediation.</p>
  • 2026-08-11T00:00:00Z
Software Engineer III (Platform Engineering)
  • San Ramon, CA
  • remote
  • Permanent / Full Time
  • 104000 - 153000 USD / Yearly
  • <p>Robert Half is seeking a <strong>Software Engineer III (Platform Engineering) </strong>to support the infrastructure, platforms, and services that power our applications and data processing environments. This role is ideal for someone who enjoys building cloud infrastructure, automating processes, improving platform reliability, and troubleshooting production issues in a fast-paced environment.</p><p><br></p><p><strong>What You&#39;ll Do</strong></p><p>Spend approximately 70% of your time building and deploying infrastructure and platform enhancements, with 30% focused on operational support and troubleshooting.</p><p>Design, build, and maintain scalable, secure, and reliable cloud-based infrastructure supporting applications, ETL/ELT processes, and platform services.</p><p>Develop and maintain infrastructure automation, monitoring solutions, CI/CD pipelines, and operational tooling.</p><p>Build and manage AWS resources, including EC2 instances and related cloud services.</p><p>Support platform services in production environments, troubleshoot outages, and drive root-cause analysis and resolution efforts.</p><p>Manage critical incidents, including coordinating with external vendors such as Microsoft to resolve high-priority production issues.</p><p>Implement Infrastructure-as-Code (IaC) solutions and improve platform reliability, scalability, and performance.</p><p>Collaborate with development, security, and operations teams to support application delivery and platform stability.</p><p>Conduct code reviews, mentor junior engineers, and promote engineering best practices.</p><p>Participate in an on-call rotation (approximately every three weeks).</p>
  • 2026-08-15T00:00:00Z
Agent Platform Engineer
  • Menlo Park, CA
  • onsite
  • Permanent / Full Time
  • 200000 - 300000 USD / Yearly
  • <p><strong>Agent Platform Engineer</strong></p><p><br></p><p><strong>Company Overview</strong></p><p>A leading artificial intelligence and advanced analytics organization is seeking an Agent Platform Engineer to help develop next-generation AI orchestration and decision-support platforms. Based in Los Angeles, California, the company specializes in integrating complex data sources into scalable intelligence solutions that support mission-critical operations, advanced analytics, and organizational workflows. This is an opportunity to work on cutting-edge AI technologies within a highly collaborative and innovative engineering environment.</p><p><br></p><p><strong>Role Summary</strong></p><p>The Agent Platform Engineer will design, build, and optimize enterprise-grade AI platforms that connect structured and unstructured data, knowledge graphs, retrieval systems, and intelligent workflows. This hands-on role focuses on developing production-ready AI applications, agent frameworks, workflow orchestration systems, retrieval pipelines, and model-serving infrastructure. The ideal candidate combines strong software engineering fundamentals with expertise in AI systems, distributed architectures, and scalable platform development.</p><p><br></p><p><strong>Key Responsibilities</strong></p><ul><li>Design and implement intelligent workflow systems for research, analysis, automation, and decision-support use cases.</li><li>Develop scalable retrieval, search, and knowledge management capabilities across structured and unstructured datasets.</li><li>Build backend services, APIs, and platform capabilities supporting AI-driven applications.</li><li>Create resilient workflow orchestration patterns for long-running, distributed, and fault-tolerant processes.</li><li>Develop authorization, permission management, auditing, and governance mechanisms within AI platforms.</li><li>Optimize retrieval, inference, and execution performance across diverse deployment environments.</li><li>Fine-tune, evaluate, and integrate open-source AI models to support specialized business workflows.</li><li>Implement model routing, inference optimization, caching, and fallback strategies.</li><li>Build evaluation frameworks that measure system accuracy, reliability, performance, and operational effectiveness.</li><li>Develop observability, monitoring, alerting, and troubleshooting capabilities across platform components.</li><li>Partner with product, data, machine learning, infrastructure, security, and customer-facing teams to deliver scalable solutions.</li><li>Contribute to technical architecture, engineering standards, and long-term platform strategy.</li></ul><p><strong>Additional Details</strong></p><ul><li>Fully onsite 5 days per week</li><li>Individual contributor role with substantial technical ownership</li><li>Opportunity to work on advanced AI, automation, retrieval, and knowledge graph technologies</li><li>Staff-level candidates may provide technical leadership, mentorship, and architecture guidance</li><li>Candidates must be authorized to work in the United States and satisfy applicable regulatory employment requirements</li></ul>
  • 2026-08-19T00:00:00Z
AI Engineer (AI Data Privacy, Security & Governance)
  • San Ramon, CA
  • remote
  • Permanent / Full Time
  • 85000 - 124000 USD / Yearly
  • <p>obert Half is seeking a Software Engineer II – AI Engineer to analyze, design, program, debug, test, implement, and support the privacy, security, and governance controls that protect generative AI technologies. This role safeguards GenAI-enabled applications, including LLM-powered workflows, RAG pipelines, plugins, skills, and autonomous agents, by embedding data protection, security guardrails, and responsible AI governance across the development lifecycle. </p><p> </p><p>This role supports SDLC documentation across all phases, with a focus on data privacy, security guardrails, governance, compliance, and risk management. It also works with users to define requirements and support applications in production. </p><p><strong>What You’ll Do</strong> </p><ul><li>Design and implement data privacy controls, including PII/PHI detection, redaction, and data minimization. </li><li>Build security guardrails, including prompt-injection defense, jailbreak prevention, and output filtering. </li><li>Establish governance for plugins, skills, and agents, including registration, approval, and lifecycle management. </li><li>Review and vet third-party plugins, skills, and agent tools for security, privacy, and compliance risks. </li><li>Define and enforce access controls, authentication, and least-privilege permissions for AI components. </li><li>Implement guardrails for autonomous agents, including action scoping, tool-use restrictions, and human-in-the-loop approval. </li><li>Apply data classification, retention, and residency policies across GenAI data flows. </li><li>Monitor AI systems for policy violations, data leakage, and anomalous plugin, skill, or agent behavior. </li><li>Maintain audit trails and logging for plugin, skill, and agent activity. </li><li>Support compliance with regulations and frameworks, including GDPR, CCPA, and the EU AI Act. </li><li>Provide Level II production support for security, privacy, and governance incidents. </li><li>Support incident response, including containment, escalation, and remediation. </li></ul>
  • 2026-08-18T00:00:00Z
Software Engineer II – AI Engineer (w/ skills in Deployment)
  • San Ramon, CA
  • remote
  • Permanent / Full Time
  • 85000 - 124000 USD / Yearly
  • <p>Robert Half is seeking a Software Engineer II – AI Engineer who will analyze, design, program, debug, test, implement, deploy, and support software enhancements and new applications using Generative AI technologies. This role contributes to the development and production deployment of GenAI-enabled applications, including LLM-powered workflows, RAG pipelines, and AI-driven user experiences. </p><p> </p><p>This role supports SDLC documentation across all phases, with a focus on deployment, evaluation, observability, safety, and monitoring. It also interacts with users to define requirements and support applications in production. </p><p><strong>What You’ll Do</strong> </p><ul><li>Develop and modify application modules, including GenAI components. </li><li>Build prompt workflows, retrieval layers, APIs, and cloud services. </li><li>Troubleshoot production issues, including latency, hallucinations, and errors. </li><li>Provide Level II production support for deployed systems. </li><li>Design components, including LLM integrations and RAG pipelines. </li><li>Implement CI/CD pipelines, containerization, and release processes. </li><li>Develop RAG pipelines with embeddings, chunking, and vector search. </li><li>Apply prompt engineering techniques, including few-shot prompting and structured outputs. </li><li>Evaluate models for accuracy, relevance, and hallucination risk. </li><li>Implement safety guardrails, including PII protection and prompt-injection defense. </li><li>Execute testing, including unit, integration, and GenAI evaluation testing. </li><li>Monitor production systems for latency, cost, usage, and errors. </li><li>Support incident management with fallback and recovery strategies. </li></ul><p><br></p>
  • 2026-08-18T00:00:00Z