About

AI systems, data infrastructure, and the craft behind reliable outputs.

This page collects the deeper profile: how I think about AI engineering, the stack I work across, and the roles that shaped the way I build.

Profile

500,000 records. Zero downtime. That's the kind of systems I build.

I engineer the infrastructure around AI: retrieval quality, document intelligence, evaluation loops, data contracts, and backend systems that make model outputs reliable enough for real users.

Sajid Shaikh with the USC Trojan horse statue
SF Bay Area, CA - USC CS alumnus

AI data infrastructure, from ingestion to intelligence

I am a USC Computer Science alumnus and AI/Data Engineer building systems across NLP text mining, RAG-ready document processing, computer vision, graph analytics, and production APIs. My focus is the hard layer between model demos and real products: clean context, trusted data, resilient services, and measurable outputs.

Data pipelines at scale 500K+ record ETL, schema validation, data quality, lineage, observability, AWS Glue, Spark, SQL
AI and RAG systems NLP text mining, RAG context optimization, multimodal AI, vision-language models, model evaluation
Backend and distributed systems Flask APIs, Java actor systems, graph analytics, CI/CD, Docker builds, zero-downtime delivery
Python Java SQL Spark Databricks AWS RAG Distributed Systems Backend Development
What's new

Learning Spanish alongside AI systems work.

I am learning Spanish as a experience indicating a small signal of curiosity, consistency, and better communication. The Ask Sajid chat now welcomes questions in English or Spanish.

Expertise

AI, data, and backend depth

My strongest work connects engineering execution with data-driven decision making: instrumenting pipelines, cleaning and joining imperfect datasets, exposing analysis through APIs, and building enough automation that teams can trust the workflow.

Data Engineering

Batch and automated pipelines, data quality checks, deduplication logic, common data models, and large SQL/PySpark workflows.

AWS Glue PySpark Databricks SQL

AI and ML Systems

Predictive models, feature engineering, fuzzy matching, similarity scoring, graph analytics, and object detection prototypes.

Pandas TensorFlow Neo4j PageRank

Backend Platforms

API design, modular services, database modeling, auth, third-party integrations, and deployment workflows for product teams.

Python Java Flask CI/CD

Languages

Python, Java, SQL, TypeScript, JavaScript, PHP

Data stack

PySpark, Databricks, Pandas, PostgreSQL, MySQL, AWS RDS, Tableau

Cloud and DevOps

AWS Lambda, Glue, Step Functions, MWAA, EC2, S3, Docker, Jenkins, Kubernetes, Git

Experience

Selected engineering journey

A focused view of the roles that shaped my AI, data, and backend engineering practice. The full resume includes additional leadership and service experience.

Feb 2026 - Present

Founding Engineer - Backend

VISIC, SF Bay Area

Architecting backend services from scratch, including REST APIs, database schemas, authentication, integrations, and deployment workflows for an early-stage product team.

Feb 2025 - May 2026

Research Assistant

USC Facilities Planning Management

  • Led a team of engineers for an NLP-based text mining system to extract and structure critical information from large-scale PDF document corpora, enabling downstream RAG context optimization and reducing manual document processing overhead at USC.
  • Directed cross-functional technical operations as Team Leader for the FMS department, maintaining an internal web platform and serving as the primary technical escalation point across stakeholders.
Jul 2025 - Aug 2025

Data Science Engineer Intern

Kognitic, Inc.

Automated structured clinical-trials data pipelines with AWS Glue, Lambda, and Step Functions, reducing manual intervention by 80% while improving duplicate detection and data consistency.

May 2025 - Aug 2025

Research Java Engineer - Distributed Systems

USC Thomas Lord Department of Computer Science

Built actor-based distributed query-engine components with Java and Apache Pekko for graph analytics workloads, including registry-driven actor discovery, streaming operators, and benchmark flows.

2021 - 2024

Business Technology Solutions Associate / Consultant

ZS Associates

Delivered Flask, SQL, PySpark, Databricks, and AWS automation for enterprise data products, optimizing 100M+ record data marts and reducing legacy analytical turnaround time by up to 90%.

Feb 2020 - May 2020

Software Developer

DeeDee Labs Private Limited

  • Designed and developed a high-performing mobile app for upper limb amputees, achieving 100% user satisfaction and boosting organizational revenue by 50%.
  • Led end-to-end development with Ionic framework, delivering production-level code within 3 months.
  • Directed a cross-functional team of 5 across product, engineering, sales, and support to launch the application with business partners.
  • Expanded hands-on expertise in SDLC, agile delivery, team leadership, Ionic, Amazon S3, PHP, MySQL, and RDBMS.
Jul 2018 - Dec 2019

Junior Java Software Developer

Jobmosis / LitmusBox

  • Worked around 40 hours per week on AI expert-system experiments and text-mining based prediction workflows.
  • Contributed to a Java-based AI prediction model deployed on AWS cloud.
  • Handled requirement understanding, low-level design, coding, integration, and testing responsibilities.
  • Demonstrated ownership, discipline, focus, and initiative while contributing to the project team.
Sajid Shaikh
connectwithsajid assistant

Ask Sajid

Hello, you are viewing Sajid's AI and Data Engineering portfolio. Ask about his technical expertise, projects, mentorship, or collaboration options. You can ask in English or Spanish.

Continue on Topmate