Production Support

Mississauga, ONTARIO, CA • Posted 9 hours ago • Updated 9 hours ago
Contract Independent
Contract Corp To Corp
Contract W2
On-site
Company Branding Image
Fitment

Dice Job Match Score™

🔗 Matching skills to job...

Job Details

Skills

  • Data Integrity
  • Continuous Improvement
  • Application Support
  • Messaging
  • Root Cause Analysis
  • Middleware
  • Database Administration
  • Incident Management
  • Production Support
  • Regulatory Affairs
  • Use Cases
  • Management Reporting
  • SLA
  • ERM
  • Risk Assessment
  • Business Rules
  • Code Refactoring
  • Data Architecture
  • System Integration Testing
  • Acceptance Testing
  • Environment Management
  • Documentation
  • Python
  • Apache Tomcat
  • Real-time
  • IT Risk Management
  • IT Risk
  • DPS
  • Data Security
  • ROOT
  • JIRA
  • Business Intelligence
  • Mentorship
  • Operational Excellence
  • Knowledge Sharing
  • Computer Science
  • Information Technology
  • Data Engineering
  • Production Engineering
  • Network Layer
  • DevOps
  • IT Operations
  • Accountability
  • Team Leadership
  • Microservices
  • Extract
  • Transform
  • Load
  • Data Modeling
  • IT Service Management
  • IT Governance
  • Communication
  • Escalation Management
  • Stakeholder Engagement
  • Leadership
  • JMS
  • Oracle
  • Data Flow
  • Tableau
  • Dashboard
  • Elasticsearch
  • Search Engines
  • MongoDB
  • Couchbase
  • NoSQL
  • Data Lake
  • Kubernetes
  • Pipeline Management
  • Continuous Integration
  • Continuous Delivery
  • GitHub
  • Bitbucket
  • Performance Monitoring
  • Java
  • Spring Framework
  • API
  • AngularJS
  • React.js
  • AppDynamics
  • Splunk
  • Kibana
  • Log Analysis
  • ServiceNow
  • Regulatory Compliance
  • CyberArk
  • Identity Management
  • Vulnerability Management
  • Collaboration
  • Management
  • SSL
  • TLS
  • Lifecycle Management
  • Financial Services
  • Reporting
  • Apache Kafka
  • Onboarding
  • Artificial Intelligence
  • Generative Artificial Intelligence (AI)
  • Workflow
  • Microsoft Certified Professional
  • Data Quality
  • Testing
  • Behavior-driven Development
  • Gherkin
  • Cucumber
  • Selenium
  • Cloud Computing
  • Migration
  • Stress Testing
  • Risk Management
  • Redis
  • Caching
  • MEAN Stack
  • Customer Service
  • Training And Development
  • SAP BASIS

Summary

Software Guidance & Assistance, Inc., (SGA), is searching for a Production Support for a right to hire assignment with one of our premier financial services clients in Mississauga, ON.

Responsibilities :
  • Responsible for the stability, resilience, data integrity, and continuous improvement of the Enterprise Risk Management (ERM) platform - a high-criticality, large-scale application comprising 17+ sub-pillars and a growing portfolio of agentic AI modules. The role sits at the intersection of production support, platform engineering, DevOps, and data platform operations, requiring deep technical capability across application support, event-driven data pipelines, and enterprise data architecture combined with strong operational discipline and cross-team coordination.
  • The successful candidate will lead an L3 engineering function, drive DevOps maturity, own data platform operations (including ERDL and the ERM Data Lake), and serve as the technical bridge between development teams, data engineering, infrastructure, and business stakeholders. This is not a passive support role - it is an active platform ownership position with accountability for production health, data quality, release governance, incident resolution, and engineering excellence.
  • Production Support & Incident Management
    • Lead L3 triage and resolution of complex production incidents across a 17+ pillar enterprise risk platform, including data ingestion failures, workflow disruptions, Kafka messaging issues, and infrastructure events.
    • Own the Problem Record (PRB) lifecycle - from triage and root cause analysis to fix coordination and post-incident documentation - in alignment with ServiceNow ITSM processes.
    • Drive the SWAT process: daily review of open problem tickets with escalation potential, ensuring senior stakeholder awareness and timely resolution.
    • Serve as the primary escalation point for L3 engineers, coordinating with development, middleware, DBA, and Tenant Ops teams as needed.
    • Lead Major Incident Management (MIM) for high-impact production events including data pipeline failures, SLA misses, and reconciliation discrepancies.
  • Data Platform Operations & Architecture
    • Own L3 production support for the Enterprise Risk Data Layer (ERDL) - the central data platform that aggregates risk data from 17+ upstream Pillar systems and Federated Limits Units (FLUs).
    • Triage and resolve data pipeline failures across Kafka, Oracle, and the ERM Data Lake - including ingestion errors, materialized view refresh failures, batch job timeouts, and data reconciliation discrepancies.
    • Perform and govern daily data quality checks: KRI count reconciliation (ERDL vs. KRI API vs. KRI/RMD Dashboard), RA count reconciliation, PRIVATE_IND NULL checks, full and incremental refresh status monitoring, and Kafka ingestion lag monitoring.
    • Maintain deep working knowledge of the ERDL reporting layer hierarchy and its authoritative use cases:
      • RMD (Risk Management Dashboard): Authoritative for business user consumption and management reporting
      • ERDL Recon View: Authoritative for reconciliation and data quality checks
      • ERDL Tableau Extract: Authoritative for historical as-of views for reconciliation purposes
    • Identify, classify, and escalate SLA Miss PRBs - distinguishing ERM application defects from upstream FLU compliance failures, and maintaining a running log of SLA miss frequency by FLU for governance escalation.
    • Support and govern the ERM Data Lake - understanding data flows, retention policies, and downstream consumer dependencies.
  • Data Modeling & Data Lineage
    • Maintain and apply working knowledge of the ERM canonical data model - understanding key entities (overlays, limits, thresholds, KRIs, risk assessments), their relationships, and how they flow through the platform.
    • Understand and document end-to-end data lineage for critical data flows: from upstream FLU systems through Kafka, into ERDL, through to downstream consumers (RMD, Tableau, Superset, Recon).
    • Apply data contract principles to triage integration failures - identifying where producer/consumer misalignments (missing fields, incorrect data types, null values, schema mismatches) are causing production defects.
    • Contribute to data architecture governance: support the validation of AI-generated data contract analyses (e.g., Devin-produced analyses), participate in Data Contract Review meetings, and help translate findings into Jira backlog items.
    • Classify production problems accurately as either Implementation Bugs (coding/configuration errors) or Policy/Process Misalignments (failures to correctly execute business rules), articulating the distinction clearly in PRB documentation.
    • Support the ERDL Refactoring initiative and related data architecture modernization efforts, including Redis infrastructure upgrades and Prod Parallel Deployment stabilization.
  • Platform & DevOps Engineering
    • Perform and coordinate DevOps functions across lower environments (DEV, SIT, UAT), including environment management, deployment sequencing, and release validation.
    • Lead or coordinate weekly production releases - owning the release bridge, executing health validation in Harness and OpenShift, and ensuring complete post-deployment documentation.
    • Drive CI/CD pipeline health across Lightspeed and GitHub - ensuring build integrity, scan compliance (Snyk, SonarQube, Checkmarx), and deployment readiness across all Pillars.
    • Manage Continuous Vulnerability Management (CVM) across the platform, coordinating remediation plans with development teams by Pillar.
    • Lead cross-pillar DevOps initiatives: Angular/React upgrades, Python version decommissions, Tomcat upgrades, Hashicorp Vault onboarding, SSL certificate lifecycle management, and GitHub migration.
  • Monitoring, Observability & Automation
    • Own and evolve the daily operational monitoring framework - Kafka ingestion health, data reconciliation, refresh status, and data quality checks.
    • Drive automation of manual monitoring activities, including automated ServiceNow incident creation for detected anomalies (e.g., via Superset or AppDynamics alerting).
    • Leverage AppDynamics, Kafka dashboards (Tableau/Superset), and OpenShift tooling for real-time platform health visibility.
    • Identify and close monitoring gaps - including proactive FID/AD group membership monitoring to prevent silent infrastructure failures.
  • Governance, Standards & Engineering Excellence
    • Enforce L3 operational standards: PRB description quality, PTASK lifecycle compliance, and Manual Touch Point (MTP) process adherence.
    • Champion Developer Manifesto compliance: README standards, GitCode ownership, branch hygiene, stale repository cleanup, and CI/CD health metrics.
    • Ensure compliance with technology risk, security, IS assessment, and regulatory standards (GIAM, EERS, CVM, CAMP, DPS data protection standards).
    • Apply the AI-Assisted Analysis and Governance framework to accelerate root cause identification and translate findings into auditable Jira backlogs.
  • Strategic & Stakeholder Leadership
    • Serve as the operational owner for new application onboarding (e.g., OMAI Overlay, Tapas, Shock Generation AI) - defining support models, escalation matrices, and runbooks.
    • Represent BAU production health and data platform status at weekly governance calls, cross-pillar bi-weekly calls, and SWAT touchpoints.
    • Partner with Architecture, Product Owners, Development Leads, Data Engineers, and Infrastructure teams to coordinate cross-pillar initiatives.
    • Mentor L3 engineers; foster a culture of operational excellence, data quality ownership, and structured knowledge sharing.
Required Skills :
  • Bachelor's degree or equivalent experience in Computer Science, Engineering, Information Technology, Data Engineering, or a related field.
  • 8+ years of technology experience, with a strong background in production engineering, L3 support, data platform operations, or platform DevOps in a large-scale enterprise environment.
  • Proven experience in a senior technical operations or engineering lead role with accountability for production stability, data pipeline health, and team coordination.
  • Demonstrated ability to triage and resolve complex, multi-system production issues across distributed microservices and data pipeline architectures.
  • Strong hands-on experience with data platform operations - including event-driven pipelines, data reconciliation processes, and multi-layer reporting architectures.
  • Working knowledge of data modeling concepts - entity relationships, canonical data models, schema evolution, and data contract principles.
  • Experience with data lineage analysis - tracing data flows from source systems through transformation layers to downstream consumers and identifying break points.
  • Strong experience with ITSM processes (ServiceNow - incident, problem, change, MTP/PRJ modules) in a formal IT governance environment.
  • Experience coordinating production releases - runbook execution, health validation, and stakeholder communication.
  • Excellent communication, escalation management, and stakeholder engagement skills - comfortable representing technical and data quality status to senior leadership.
  • Data Platform & Architecture
    • Kafka / JMS - producer/consumer health monitoring, topic-level triage, schema validation, consumer lag analysis
    • Oracle - read-level query capability; working knowledge of materialized views, batch jobs, schema structures, and data reconciliation views
    • Data lineage tooling - ability to trace and document data flows across multi-system architectures
    • Data contract principles - producer/consumer responsibilities, schema validation, null-safety, field-level contract analysis
    • Tableau / Superset - operational dashboard monitoring and data layer reconciliation
    • Elastic Search - basic operational awareness for search/index layer triage
    • MongoDB / Couchbase - operational awareness for NoSQL data stores in use across Pillars
    • Data Lake concepts - retention policies, data classification, downstream consumer patterns
  • Platform & Infrastructure
    • OpenShift / Kubernetes - pod management, health checks, container operations
    • Harness - deployment pipeline management and release validation
    • Lightspeed (LSE / Classic) - CI/CD platform management
    • GitHub / Bitbucket - repository governance, branch management, pipeline configuration
    • AppDynamics - application performance monitoring and alerting
  • Application & Integration
    • Java / Spring Boot - sufficient depth to triage application-layer issues, interpret stack traces, and understand data contract failures
    • REST APIs - API failure interpretation, connectivity validation, integration troubleshooting
    • Angular / React - basic familiarity for front-end issue triage
  • Observability & Tooling
    • AppDynamics, Splunk, ELK / Kibana - log analysis and alerting
    • SonarQube, Snyk, Checkmarx - compliance gate interpretation
    • ServiceNow - incident, problem, change, and PRJ module management
  • Security & Compliance
    • CyberArk, CISAR - FID and privileged access management
    • EEMS / EERS - entitlement management and access review processes
    • CVM / CAMP - vulnerability management and Pillar-level remediation coordination
    • Hashicorp Vault - secrets management operations
    • SSL / TLS certificate lifecycle management
Preferred Skills :
  • Experience in financial services or regulated technology environments strongly preferred.
  • Experience with ERDL or equivalent enterprise risk data layer platforms - aggregating data from multiple upstream systems into a central risk reporting layer.
  • Familiarity with Kafka schema governance and event-driven integration patterns between enterprise risk systems (limits, thresholds, KRIs, model risk).
  • Experience supporting or onboarding agentic AI or GenAI-integrated applications (e.g., Generative AI overlays, LLM-backed workflows, MCP-based architectures).
  • Experience implementing automated data quality monitoring and self-healing alerting frameworks.
  • Knowledge of FAST automation framework, contract testing, or behavior-driven development (Gherkin/Cucumber/Selenium).
  • Familiarity with cloud modernization (Cloud @ large firm / Type A migration) and container-native platform evolution.
  • Exposure to enterprise risk management concepts - 1LOD/2LOD governance, stress testing (CCAR/QMMF), model risk management, or limits and thresholds frameworks.
  • Experience with Redis infrastructure for caching and entitlement stability.
SGA is a technology and resource solutions provider driven to stand out. We are a women-owned business. Our mission: to solve big IT problems with a more personal, boutique approach. Each year, we match consultants like you to more than 1,000 engagements. When we say let's work better together, we mean it. You'll join a diverse team built on these core values: customer service, employee development, and quality and integrity in everything we do. Be yourself, love what you do and find your passion at work. Please find us at .

SGA is an Equal Opportunity Employer and does not discriminate on the basis of Race, Color, Sex, Sexual Orientation, Gender Identity, Religion, National Origin, Disability, Veteran Status, Age, Marital Status, Pregnancy, Genetic Information, or Other Legally Protected Status. We are committed to providing access, equal opportunity, and reasonable accommodation for individuals with disabilities in employment, and our services, programs, and activities. Please visit our company to request an accommodation or assistance regarding our policy
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: sgainc
  • Position Id: 26-01932
  • Posted 9 hours ago

Company Info

About Software Guidance & Assistance

Founded in 1981, SGA is a technology and resource solutions provider with a national footprint and headquartered in the shadow of Wall Street. We’re a certified women-owned business. We provide contingent staffing, direct placement, and professional and managed services to transform businesses and evolve careers. We’re small enough to tailor our services to each client and big enough to deliver for some of the world’s largest employers. Our professionals are experts in areas such as IT, finance, accounting, risk, and clinical.

SGA provides contingent staffing, direct placement, and professional and managed services nationwide for Fortune 500 companies, mid-size businesses and select startups.

Our core skillsets include all areas of technology – business & data analysis, cyber & network security, database administration, development & architecture, infrastructure, program & project management, quality assurance & testing. We also deliver talent across professional business functions such as finance, accounting, risk, and clinical.

Our Professional & Managed Services team delivers IT projects through onshore, offshore and hybrid delivery models. We develop software products, modernize applications, add features, and integrate and maintain systems. Our scope covers, among others, complex application suites, data management and visualizations, machine learning and mobile applications.

About_Company_OneAbout_Company_Two
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Mississauga, Ontario

Today

Contract, Third Party

Mississauga, Ontario

Today

Contract, Third Party

Mississauga, Ontario

Today

Contract, Third Party

Holmdel, New Jersey

Today

Contract

USD 50.00 - 58.00 per hour

Search all similar jobs