Vijay Sojitra

Full-Stack AI/ML Engineer & Applied Statistician

Agentic AI Engineering • LLMOps • MLOps • Data Science • Data Engineering • GraphRAG • AI Web Dev

Columbus, Ohio · vjsojitra8@gmail.com

Builds the complete path from governed data pipelines and statistical modeling through experimentation, model and LLM deployment, evaluation, monitoring, APIs, and decision-support products.

10 years Enterprise AI Systems Design, Machine Learning, Data Engineering & Science and Business Analytics
Dual Master’s Degrees Statistics and Statistical Data Science

SELECTED WORK

Enterprise AI/ML Systems

Large-Scale Data Engineering

Package Network Fingerprint & Temperature Intelligence

Large-scale Azure Databricks/PySpark pipeline that derives package-network fingerprint and temperature intelligence from multi-billion-row operational source data.

  • Native PySpark transformations with partition-aware joins, window operations, column pruning, and selective caching—no Python UDFs
  • Approximately 4.5 billion source rows producing ~200 million Fingerprint records and ~22 million Temperature outputs with a normalized score from cold 0 to hot 1
  • Row-count reconciliation, duplicate detection, and null/range validation

Enterprise GenAI

Governed Self-Service Operations Intelligence

Employee-facing AI system that routes natural-language requests between structured-data querying and cited document retrieval under enterprise governance controls.

  • Natural-language-to-SQL with schema grounding, approved patterns, validation, plausibility checks, and restricted execution
  • Azure OpenAI and authenticated enterprise APIs with Azure AI Search hybrid retrieval, semantic chunking, reranking, citations, and answerability thresholds
  • MLflow evaluation for retrieval precision, answer quality, hallucination rate, latency, and failure modes, with human review where operational risk required it

Document AI

DOT Document Intelligence & Decision Support

Document intelligence workflow for DOT decision support, combining extraction, reviewer queues, and LLM-assisted evidence validation.

  • AWS Textract, custom neural networks, and Python normalization for images and Word documents, with LLM-assisted evidence validation
  • Reviewer queues with lineage, confidence indicators, and exception handling
  • Source project records report 70% faster processing at approximately 90% extraction accuracy

Applied Machine Learning

Predictive Maintenance Operating Model

Closed-loop predictive maintenance spanning telemetry, model alerts, and technician-facing service recommendations across manufacturing sites.

  • ARIMA, XGBoost/LightGBM, anomaly detection, and computer vision applied to distinct maintenance problems
  • Event-driven alerts with technician-facing service recommendations
  • Azure AI/IoT capability scaled across 50+ manufacturing sites in a closed loop from telemetry and validation through alerting and technician feedback

Independent Engineering

These projects demonstrate independently built systems across product engineering, multi-agent orchestration, agent operations, LLM reliability, retrieval, machine learning, MLOps, and vertical software development.

Multi-Agent Systems Engineering

Hierarchical Multi-Agent Orchestration Framework

A controlled system for dividing complex software work among specialized AI agents without letting unvalidated or failed outputs silently affect later stages.

  • Top-level orchestrator decomposes requests and delegates bounded tasks to frontend, backend, QA, performance, review, and validation agents
  • Shared execution state with dependency-aware handoffs, task identities, artifact lineage, and structured workflow events
  • Success/failure gates, fail-closed propagation, human escalation checkpoints, and inspectable Markdown and JSON run reports

Agentic AI Infrastructure & Operations

Agentic AI Operating System — OpenClaw, Hermes and Vault Memory

A persistent personal AI environment connecting communication channels, specialized agents, model providers, project memory, operational controls, and recovery.

  • OpenClaw communication and routing with Hermes execution profiles and controlled delegation through Discord and Telegram
  • Provider-selection and fallback patterns to reduce dependence on one local model server or hosted inference provider
  • Obsidian Vault-as-Memory with Git history, source-of-truth records, encrypted backups, integrity checks, health monitoring, and operator controls

Vertical SaaS & Product Engineering

PriceBook Forge — Offline Pricing Software for Home-Service Contractors

A browser-based tool that helps small HVAC, plumbing, and electrical businesses calculate sustainable flat-rate prices and produce a branded customer price book.

  • Calculates break-even and target retail hourly rates from labor, overhead, utilization, and margin assumptions with transparent math
  • Editable trade-specific repair tasks that automatically generate Good-Better-Best customer pricing
  • Branded printable output with financial assumptions kept in-browser—never transmitted to an application server

LLM Engineering & Reliability

Prompt Compiler — Validated, Self-Correcting LLM Pipelines

A framework that turns ambiguous LLM requests into structured execution contracts and rejects or corrects outputs that fail defined requirements.

  • Converts requests into roles, objectives, context, constraints, processes, output schemas, acceptance criteria, and stop conditions
  • Pydantic or equivalent schema validation plus critic review to catch structural and semantic failures before acceptance
  • Bounded correction loops that feed validation errors back and record execution evidence for testing, debugging, and audit

RAG & AI Application Engineering

IT Support Intelligence — Hybrid Retrieval and RAG Platform

A support application that finds relevant historical IT resolutions with keyword and semantic search, then optionally generates a proposed response.

  • BM25 keyword search and FAISS vector similarity fused with Reciprocal Rank Fusion so results are not tied to one method
  • Sentence-transformer embeddings, FastAPI endpoints, Gradio interface, caching, structured logging, rate controls, and fallbacks
  • Approximately 7,000 synthetic support tickets plus retrieval and API tests for indexing, search, and service behavior

Machine Learning & MLOps

Default-Risk ML Service — Training, Tracking, CI/CD and Serverless Inference

A machine-learning service that trains a credit-card default-risk model, tracks experiments, packages inference behind an API, and defines automated cloud deployment.

  • TensorFlow and scikit-learn training with MLflow experiment tracking and versioned model artifacts
  • FastAPI prediction service in Docker with configurable model and preprocessing artifact loading
  • GitHub Actions for testing, image building, deployment workflow, and Google Cloud operational logging

EXPERIENCE

AI Engineer

FedEx via TEKsystems
  • Contributed to billion-row package-network intelligence pipelines in Azure Databricks and native PySpark, including Fingerprint and Temperature transforms with validation gates.
  • Built and implemented governed natural-language querying and cited document retrieval for employee-facing operations intelligence, with schema grounding, restricted execution, and human review where risk required it.
October 2025 – March 2026

Senior Data Scientist / AI Solutions Architect

HNTB Corporation
  • DOT document intelligence and decision support using AWS Textract, custom neural networks, reviewer queues, and LLM-assisted evidence validation.
  • Reusable ML release and monitoring platform with standardized health monitoring across 15+ models and an approximately 2 TB/day anomaly-detection pipeline.
  • Governed RAG assistant and centralized AI Gateway evaluated on answer quality, citation support, latency, and cost; reported 60% reduction in deployment time while leading delivery teams of up to 8–12.
August 2023 – March 2025

Data Scientist

HCL Global, Capital One engagement
  • Built PySpark ETL on AWS EMR/Hadoop processing billions of rows monthly, with Snowflake, Python, SQL, and Salesforce API integration.
  • Implemented Airflow batch orchestration and Kafka streaming inputs for model-development workflows.
  • Developed XGBoost, LightGBM, and TensorFlow models; transformer semantic-similarity features reducing false positives by 22%.
  • Tracked experiments with MLflow across 200+ recorded model versions and contributed Docker/Kubernetes/SageMaker deployment into an existing serving environment.
July 2022 – June 2023

Data Scientist

Home Depot via Ugam Solutions
  • Retail product-review and merchandising signal analysis with reproducible feature and scoring datasets.
  • A/B-testing frameworks and Hugging Face multilingual review processing across eight languages (~100,000 product reviews per month).
  • BigQuery scoring/reporting automation; reported 30% lower processing time.
February 2022 – June 2022

Data Scientist

Crown Equipment Corporation
  • Connected-equipment analytics foundation for predictive maintenance and service recommendations.
  • Azure AI/IoT program scaled across 50+ manufacturing sites with technician feedback and operational adoption loops.
  • Source records report 60% lower downtime and approximately $1 million in attributed additional service revenue.
August 2019 – January 2022

Earlier Analytics & Applied Data Science

Ugam Solutions · HDFC Bank · PRM Fincon · Evolvalytics
  • Progression from SAS/SPSS/Excel automation to Python, SQL, and R pipelines.
  • Credit-risk modeling and statistical testing; forecasting, segmentation, and inventory decision support.
2011–2017

SKILLS

GenAI & Agentic Systems

RAG and hybrid retrieval · Semantic chunking and reranking · Citation enforcement and hallucination controls · LangChain · LangGraph · LlamaIndex · Tool/function calling · Human-in-the-loop workflows · Orchestrator/executor patterns

MLOps / LLMOps

MLflow · Docker · Kubernetes · Jenkins · GitHub Actions · Experiment tracking · Model and prompt evaluation · Data/model validation gates · Drift and performance monitoring · Audit logging and rollback planning

Data, Cloud & Platforms

Databricks · Snowflake · AWS EMR/S3/Lambda/SageMaker/Textract · Azure OpenAI · Azure AI Search · Azure ML · Azure Data Factory · GCP/BigQuery · Kafka and Airflow

ML & Applied Statistics

scikit-learn · XGBoost · LightGBM · TensorFlow · PyTorch · Forecasting · Anomaly detection · NLP · A/B testing · Explainability and bias checks

Core / Primary

Python · SQL · PySpark/Spark · Statistical modeling · Feature engineering · ETL/ELT architecture · Experimentation and model validation


LEADERSHIP & ENGINEERING APPROACH

Technical Leadership

  • Led delivery teams of up to 8–12
  • Mentored junior data scientists, AI engineers, and interns
  • Supported sprint planning, architecture discussions, and technical standards

Enterprise Translation

  • Converted ambiguous operational problems into measurable systems
  • Presented architecture, risk, and platform trade-offs to technical and executive stakeholders
  • Coordinated data, ML, platform, security, and governance teams

Engineering Principles

  • Reproducibility
  • Explicit validation
  • Observability
  • Failure analysis
  • Security and governance
  • Human review where risk requires it

EDUCATION

Florida State University

Master of Science in Statistical Data Science
August 2017 – May 2019

Sardar Patel University

Master of Science in Statistics
August 2012 – May 2014

CONTACT

Open to discussing AI engineering, machine learning, statistical modeling, data/ML platforms and AI automation work.

vjsojitra8@gmail.com