Software Developer / Engineer - Philadelphia, PA (Locals Only)

Philadelphia, PA, US • Posted 3 days ago • Updated 3 days ago
Contract W2
Contract Independent
Contract Corp To Corp
12 Months
No Travel Required
On-site
$45 - $70/hr
Fitment

Dice Job Match Score™

👾 Reticulating splines...

Job Details

Skills

  • Python
  • LLM
  • RAG

Summary

Software Developer / Engineer
Location: Philadelphia, PA 
Work Schedule: Hybrid 3 days on site, 2 remote


Position Overview

We are seeking a Software Developer / Engineer to help design and implement an on-premises Large Language Model (LLM) platform with Retrieval-Augmented Generation (RAG) capabilities. This role will focus on deploying open-source AI models, integrating vector databases, and building secure, enterprise-grade AI solutions in a private environment.

This is an excellent opportunity for a developer with hands-on experience in modern AI technologies who enjoys building scalable, high-performance systems.

Responsibilities

  • Deploy and optimize open-source large language models (LLMs) such as Meta Llama 3 and Mistral/Mixtral in on-premises or private environments.
  • Develop Python-based applications for LLM inference, prompt engineering, and model integration.
  • Optimize CPU-based model inference through quantization and performance tuning.
  • Design and implement Retrieval-Augmented Generation (RAG) (RAG) pipelines.
  • Configure and manage open-source vector databases such as Qdrant, Chroma, Milvus, or pgvector.
  • Generate and manage embeddings while implementing metadata filtering strategies.
  • Support enterprise security requirements, including air-gapped deployments, access controls, data privacy, and audit logging.
  • Produce technical documentation, deployment guidance, and knowledge transfer materials for internal teams.
  • Build a working prototype integrating an LLM, vector database, and RAG architecture.

Required Qualifications

  • Professional experience deploying open-source LLMs (e.g., Meta Llama 3, Mistral/Mixtral) in on-premises or private environments.
  • Strong Python development experience.
  • Hands-on experience with LLM inference, prompt engineering, and AI application integration.
  • Experience optimizing CPU-based inference through model quantization and performance tuning.
  • Experience with vector databases such as Qdrant, Chroma, Milvus, or pgvector.
  • Proven experience implementing Retrieval-Augmented Generation (RAG) solutions.
  • Understanding of enterprise security, data privacy, air-gapped environments, access controls, and audit logging.

Preferred Qualifications

  • Experience with LangChain or LlamaIndex.
  • Familiarity with Docker and Kubernetes.
  • Experience with inference frameworks such as vLLM, llama.cpp, or Hugging Face Transformers.
  • Experience with Rust, Go, or C++.
  • Previous experience working in enterprise or regulated environments.

Deliverables

  • Reference architecture and deployment guidance.
  • Working prototype integrating an LLM, vector database, and RAG solution.
  • Technical documentation and knowledge transfer to internal teams.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10513292
  • Position Id: 73215-12895-1784833053
  • Posted 3 days ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Hybrid in Philadelphia, Pennsylvania

Yesterday

Easy Apply

Contract

Depends on Experience

Hybrid in Philadelphia, Pennsylvania

2d ago

Easy Apply

Third Party, Contract

Depends on Experience

Hybrid in Philadelphia, Pennsylvania

4d ago

Easy Apply

Contract, Third Party

Depends on Experience

Remote or Philadelphia, Pennsylvania

Today

Easy Apply

Contract

$DOE

Search all similar jobs