Senior Software Test Engineer

Remote in Austin, TX, US • Posted 6 hours ago • Updated 6 hours ago
Full Time
On-site
USD 165,000.00 per year
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Valuation
  • Network
  • Health Care
  • Collaboration
  • Health Insurance
  • Recruiting
  • Test Plans
  • Writing
  • Management
  • IC
  • Integrated Circuit
  • Internal Communications
  • Continuous Integration
  • System Integration Testing
  • Workflow
  • Dashboard
  • ROOT
  • CHAOS
  • Quality Assurance
  • Exploratory Testing
  • Evaluation
  • Regression Analysis
  • Debugging
  • Forensics
  • Python
  • TypeScript
  • Fluency
  • IT Operations
  • Value Engineering
  • Sprint
  • Communication
  • Artificial Intelligence
  • LangSmith
  • Project Management
  • Performance Management
  • Preventive Maintenance
  • Scratch
  • ARM
  • Testing
  • Pharmacy

Summary

About Curative

Curative is building the future of health insurance with a first-of-its-kind employer-based plan designed to remove financial barriers and make care truly accessible: one monthly premium with $0 copays and $0 deductibles*. Backed by our recent $150M in Series B funding and valuation at $1.275B, Curative is scaling rapidly and investing in AI-powered service, deeper member engagement, and a smart network designed for today's workforce.

Our north star guides everything we do: healthcare only works when people can actually use it. That belief drives every decision we make: from how we design our plan, support our members, to how we collaborate as a team.

If you want to do meaningful work with a team that moves fast, experiments boldly, and cares deeply, Curative is the place to do it. We're growing fast and looking for teammates who want to help transform health insurance for the better.

Summary
We are hiring a Senior Software Test Engineer to help us deliver quickly with the right level of quality. If you're picturing test plans, sign-off gates, and sprint ceremonies, this is not that job. At Curative, every engineer owns what they ship. Your job is to make that possible at speed. A big part of that is our AI agents, which do real operational work in production. We're investing in the evaluation layer to match: eval suites, regression harnesses, and systematic measurement. But the role is not only infrastructure. You'll work directly with business teams, become a domain expert in corners of our business, test the way real users work, and catch the gap between what was asked for and what was shipped. Some days you're writing an eval harness; some days you're hunting a bug an operations teammate can feel but can't pin down; some days you're shaping what gets built next. You'll advise on quality concerns, and you'll jump in and fix things yourself: same day, not next sprint. If ambiguity sounds stressful, this is the wrong role. If moving between code, product, and people in the same afternoon sounds like the best part of the job, keep reading.

How we build
At Curative, AI writes most of the code. Engineers direct it, using agentic AI coding tools as the primary development surface: setting context, making the decisions the AI cannot, and keeping the bar high on what ships. This is not a role for someone who wants to hand-roll every line, nor for someone who will accept whatever the AI produces. We want the engineer in between: fundamentals strong enough to catch a wrong answer fast, discipline to review every diff, and ambition to drive several times the output of a traditional IC.

What you'll own

  • The eval platform. Harnesses, datasets, graders, and CI integration that let teams measure agent behavior: task success, tool-use correctness, drift, latency, cost. You'll build it and make it the thing teams reach for because it's useful, not because a process says so.
  • Hands-on, exploratory testing. You test the way members, operations teams, and providers actually use the product, end to end, weird paths included, and find what breaks before the real world does.
  • Validating we built the right thing. You sit with business stakeholders, understand the real workflow, and catch the gap between what was asked for and what was shipped.
  • Domain expertise. You go deep enough on our business that teams pull you in because you understand the workflow, not just the code.
  • Production signal and bug hunting. Traces, failure taxonomies, and dashboards that turn "the agent seems off" into a specific finding. In incidents you reproduce the failure, find the root cause, and fix it or hand off a diagnosis so precise the fix is obvious.
  • Product judgment when needed. Sometimes there's no PM in the room. You can write the requirement, propose the behavior, make the call, and hand it back gracefully when the right owner shows up.
  • Prioritization under chaos. Finite hours, wide surface. You decide where effort goes first, which agents carry the most risk, and which failures are expensive versus merely embarrassing, and you defend that call.

What we're looking for

Foundational skills

  • 5+ years in software quality, testing, or engineering with real range. You've written automation and done serious exploratory testing, moving between them as the problem demands.
  • Real fluency with LLM evaluation. Golden datasets, LLM-as-judge, programmatic graders, regression suites for prompts and agent loops. You can point to evals that caught real problems.
  • Hands-on experience with AI agents and the ways they fail: silent drift, tool misuse, compounding errors.
  • A tester's instincts. You find the bug nobody else does, and you have the track record to prove it.
  • Debugging depth. Logs, traces, queries, production forensics. You find the bug and can fix it yourself in Python or TypeScript.
  • Fluency with business stakeholders. You can run a session with a non-technical operations lead, extract what they actually need, and leave them feeling heard.
  • Comfort with ambiguity and speed. You've been effective without mature processes. You don't need a ticket or sprint boundary to start.
  • Pragmatic quality judgment. You can say "this risk is acceptable, ship it"as often as "this one isn't"
  • Sharp written communication. Your findings change what people do.

AI-first working style

  • Claude Code, Cursor, or equivalent is already your primary development tool.
  • You use AI to accelerate the quality work itself (eval cases, graders, triage) and know where AI judgment ends and yours begins.
  • You review every AI-generated diff. You do not merge on vibes.

Strongly preferred

  • Stood up eval or observability infrastructure from zero, and it stuck.
  • LLM observability tooling: LangSmith, Braintrust, Arize, or homegrown.
  • Defined what an agent may do unsupervised and built the guardrails.
  • Worn a product hat: requirements, behavior, de facto PM for a surface.
  • Built deep domain expertise from scratch in a complex operational business.

Why this role
Most testing roles keep you at arm's length from the business and the product. This one doesn't. Every engineer here owns what they ship. You give the whole team the tools, testing, and domain insight to move fast and get it right. Eval infrastructure, hands-on testing, real relationships with the business, and sometimes the product call itself: if you want that range with real production stakes and no bureaucracy in the way, this is that role.

Perks & Benefits

  • Curative Health Plan (100% employer-covered medical premiums for you and 50% coverage for dependents on the base plan.)

    • $0 copays and $0 deductibles (with completion of our Baseline Visit )
    • Preventive and primary care built in
    • Mental health support
    • One-on-one care navigation
    • Chronic condition programs (diabetes, weight, hypertension)
    • Maternity and family planning support
    • 24/7/365 Curative Telehealth
    • Pharmacy benefits
  • Comprehensive dental and vision coverage
  • Employer-provided life and disability coverage with additional supplemental options
  • Flexible spending accounts
  • Generous PTO policy plus 11 paid annual company holidays
  • 401K for full-time employees
  • Generous Up to 8-12 weeks paid parental leave, based on role eligibility.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: RTX1e080f
  • Position Id: 3a01114578610e01839a3aa5ba3d8fae
  • Posted 6 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Full-time

USD 240,000.00 - 265,000.00 per year

Remote

Today

Full-time

USD 146,200.00 - 190,000.00 per year

Remote

11d ago

Easy Apply

Full-time, Third Party

Depends on Experience

No location provided

Today

Full-time

USD 89,200.00 - 209,500.00 per year

Search all similar jobs