Role: Deep Learning Consultant
Location: Santa Clara, CA
Onsite (5 days/week)
Duration: 12 Months
Exp: 12+yrs
Training Large Language Models
Good keywords to look for;
- Distributed training: FSDP (Fully Sharded Data Parallel), DeepSpeed, ZeRO (Zero Redundancy Optimizer), Megatron-LM (Megatron-Language Model), tensor/pipeline/data parallelism, multi-node, NCCL (Nvidia Collective Communoications Library)
• Fine-tuning internals: LoRA, QLoRA (Quantized LoRA), PEFT (Parameter Efficient Fine Tuning), SFT, full fine-tuning, instruction tuning, pretraining from scratch
• Post-training/alignment: RLHF, DPO, PPO, reward modeling
• The mechanics: gradient checkpointing/accumulation, mixed precision (bf16/fp16), learning-rate warmup/scheduling, loss curves, convergence, checkpointing
• Right-altitude tooling: raw PyTorch, Hugging Face Transformers/Accelerate, JAX/Flax — plus H100/A100 clusters, CUDA
Validating LLMs
Good Keywords:
Eval harnesses (lm-evaluation-harness), perplexity, held-out/custom eval sets
• Benchmarks: MMLU, HELM, HellaSwag
• Hallucination measurement, red-teaming, human eval, LLM-as-judge
Client Comments:
Questions specific related to training and validating large language model is what we can focus on.
It is important for us to validate if they claim on training language models, they were hands on and have the good knowledge on the concepts rather than just running via prebuilt libraries or services.