Our client’s unstructured estate — thousands of SharePoint sites of varying ownership and activity, department drives, OneDrive, plus structured stores — cannot be mapped by interviews alone. This role stands up AI-enabled discovery and classification to find personal information, tag it to our client’s record types and populate the customer and associate data maps, with a human validation step at every stage. Phase 1 runs on a bounded dataset (and on our client’s existing Microsoft capability if external tooling has not yet cleared security review); Phase 2 scales across nine functions; Phases 3–4 hand over classification rules and re-scan to prove evidence traceability.
Nothing you build touches our client production customer or employee data until it has passed our client’s AI vendor/tool risk assessment and data processing and security review — you will help prepare that submission in week one. Output is never accepted on model confidence alone; the Data Engineer independently validates it.
What you will do, by phase
Phase / commitment | | What you will do | |
Phase 1 — Mobilize & Assess (wk 1–6) Full | | ● Prepare and submit the candidate tooling package (architecture, data flows, processing location, retention of scan results, access model) to the client’s InfoSec team for the AI vendor/tool risk assessment ● Deploy discovery tooling inside our client’s environment; run a bounded-dataset classification pilot (recommended: People & HR record types) and measure precision/recall against human review ● Define sensitive information types / classifiers for our client’s record taxonomy, including SPI and ePHI flags | |
Phase 2 — Execute in Waves (wk 3–16) Full | | ● Run discovery and classification across the remaining nine functions in waves aligned to the Compliance Business Analyst’s schedule ● Populate data maps: personal data elements, collection source, purpose, internal recipients/flows, external recipients, system location, record type, retention trigger/period ● Present AI output alongside independently validated findings; maintain a sampling-based validation log and error analysis ● Automate governance documentation where it accelerates the work (workbook population, lineage narratives) with review gates | |
Phase 3 — Governance & Operating Model (wk 5–18) Partial | | ● Hand over classification rules, classifiers and scan configurations with documentation so the client can operate them; define re-scan cadence for the health-check process | |
Phase 4 — Evidence & Audit Support (wk 17–28) Partial | | ● Run re-scan validation to confirm disposition took effect and no untracked copies remain; produce evidence traceability from scan → inventory → disposition record ● Answer auditor questions on how the inventory was produced and validated | |
Deliverables you own or co-own
● Enterprise Data Inventory and Repository Register — AI discovery output presented alongside independently validated findings (co-owned with Data Engineer)
● Customer and Associate Data Maps — populated for every record type carrying personal information (co-owned with Compliance Business Analyst)
● Classification rule handover pack and re-scan evidence (Phases 3–4)
● AI tooling risk-assessment submission and approval record
Must-have qualifications
● 5+ years applied machine learning, data science or automation engineering with hands-on data classification / sensitive-data discovery at enterprise scale
● Production experience with at least one discovery/classification platform: Microsoft Purview (sensitive info types, trainable classifiers, DSPM), BigID, Varonis, OneTrust or equivalent
● Designed human-in-the-loop validation with measurable quality (precision/recall, sampling plans, error analysis) and can explain results to non-technical reviewers
● Python; Microsoft Graph / SharePoint APIs or equivalent for estate-scale scanning; secure handling of results inside a client tenant
● Has prepared or contributed to an enterprise AI/security review (architecture, data flow, model/data residency) and shipped under its constraints
● Understands personal information categories under CCPA/CPRA, SPI and ePHI well enough to design classifiers for them
Strongly preferred
● Azure OpenAI / LLM-based document classification pipelines with evaluation harnesses; agentic workflows for governance documentation with review gates
● BigID or Purview certifications; NIST AI RMF familiarity
● Experience classifying legacy formats (mainframe extracts, COBOL copybooks, call-recording transcripts)
● Prior privacy or records program where classification output fed retention labels or a data map
Tools and platforms
Microsoft Purview Information Protection / DSPM, BigID, SharePoint Online and Microsoft Graph, Azure services as approved by the client, Python, Excel function workbooks, Clarity PPM.