Job DescriptionWe are looking for seasoned and talented developers to join the Observability & Supervision team within Oracle Cloud Infrastructure (OCI). Our team is part of Autonomous OCI, a new service platform powering OCI services that encompasses low-level execution runtime, application management and high-level workflows for change management. Our mission is to help OCI developers concentrate on building their services while we ensure they run efficiently and reliably at cloud scale.
Observability & Supervision is how AOCI sees and heals itself. We build the systems that collect and query metrics and logs across the fleet, evaluate the health of every workload in real time, and decide what to do when something goes wrong. Our goal is an autonomous platform: one that detects, diagnoses and recovers from failures on its own, scales itself ahead of demand, and escalates to humans only when it truly needs to. This means building low-latency, highly available distributed systems where milliseconds and correctness matter, and where the platform is only as reliable as the component watching over it.
These are exciting times, and our team is still new and growing. This is your opportunity to build ambitious new initiatives with broad impact across OCI. We need you to challenge existing engineering assumptions and boundaries, bring your expertise in highly performant, reliable software engineering and help us bring OCI to the next level. We are looking for engineers who are self-motivated, passionate about solving complex software challenges, and able to dive deep into systems to understand and improve them. You should value simplicity and scale, work comfortably in a collaborative, agile environment, and be excited to learn.
What You Will Do Design and build the supervision control loop of the Kilt platform: health evaluation, automated recovery, escalation policies and self-healing workflows.
Build scalable telemetry pipelines for metrics and logs, including ingestion, storage, querying and alerting, operating at OCI fleet scale.
Develop platform triggers that drive autonomous actions such as recovery and autoscaling from live platform state, with tight end-to-end latency budgets.
Write performance-critical, memory-safe services in Rust alongside Java and Go components.
Engineer for high availability: replication, partitioning, failover and graceful degradation, validated through load, performance and chaos testing.
Partner with runtime, compute and service teams across OCI to make observability and recovery a built-in property of every service running on Kilt.
ResponsibilitiesBasic Qualifications BS or MS degree in Computer Science or relevant technical field involving coding, or equivalent practical experience
4+ years of full-time professional experience in software development
Demonstrated ability to write great code using Rust, Java, GoLang, or similar languages
Proven ability to deliver products and experience with the full software development lifecycle
Experience working on large-scale, highly distributed services infrastructure
Experience working in an operational environment with mission-critical tier-one livesite servicing
Systematic problem-solving approach, strong communication skills, a sense of ownership, and drive
Experience designing architectures that demonstrate deep technical depth in one area, or span many products, to enable high availability, scalability, market-leading features and flexibility to meet future business demands
Preferred Qualifications Experience as technical lead on a large-scale cloud service
Production experience with Rust, particularly for low-latency or high-throughput systems
Experience building observability systems: metrics, logging, tracing, time-series databases, query engines or alerting platforms
Experience building automated remediation, self-healing or autoscaling systems, and designing control loops that act safely without human intervention
Knowledge of distributed systems fundamentals: consensus, replication, gossip protocols, partitioning and failure detection
Hands-on experience developing and maintaining services on a public cloud platform (e.g., AWS, Azure, Oracle)
Experience working on Kubernetes
Knowledge of Infrastructure as Code (IaC) languages, preferably Terraform
Strong knowledge of databases (SQL and NoSQL), including embedded storage engines
Strong knowledge of Computer Networking (OSI layers, HTTP, DNS, TCP/IP, DHCP, Routers, Gateways, Subnets, etc.)
Knowledge of Linux internals, Linux/Unix troubleshooting skills
Familiarity with host virtualization technologies (KVM, Containers, Docker, etc.)
Able to effectively communicate technical ideas verbally and in writing (technical proposals, design specs, architecture diagrams and presentations)
- Experience with hiring, mentorship and raising the talent bar
QualificationsDisclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.Range and benefit information provided in this posting are specific to the stated locations onlyUS: Hiring Range in USD from: $79,200 to $209,500 per annum. May be eligible for bonus and equity.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
As part of Oracle's onboarding process and consistent with applicable law, US-based employees are required to complete identity verification, which involves the collection and processing of their biometric information. Accommodations to this requirement may be granted following an individualized assessment.
About UsOnly Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing or by calling 1- in the United States.
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.