Venkatesh Sridharan

Venkatesh Sridharan

Data Engineer · Toronto, ON

Canadian Permanent Resident – no sponsorship required

I'm a data engineer with 7+ years building and owning production data platforms for financial-intelligence and analytics products. I work hands-on across the AWS data stack (S3, Glue, Redshift, Athena, Step Functions), Airflow orchestration, and dbt transformation modeling, and I design LLM-orchestrated data pipelines – entity extraction, classification, and enrichment workflows built on the OpenAI and Gemini APIs.

My track record is taking ETL systems from bottleneck to reliable at scale, and growing team capability through hands-on mentoring.

Experience

Data Engineer

PrivCo

  • Architected and owned ETL data pipelines and dbt transformation models across AWS (S3, Glue, Step Functions, Redshift), cutting processing latency 30% and raising the reliability of datasets feeding ML training and downstream analytics.
  • Designed and led development of an AI-driven company-profile ingestion pipeline (n8n, Gemini LLM, Python) that transforms vendor datasets into structured company profiles – descriptions, NAICS/PICS classification, keyword metadata – automating onboarding of 10k+ companies per batch and materially cutting manual analyst effort. Read the case study
  • Built and deployed a real-time news classification system ingesting 1K+ articles/day from 20+ sources (Python, Airflow, AWS), auto-tagging IPO, M&A, Funding, and Delisting events and eliminating 100+ analyst hours/week.
  • Built LLM-powered entity extraction and enrichment pipelines (OpenAI API, n8n), improving financial-event tagging accuracy 25% within 6 months and expanding automated coverage of company-level financial signals.
  • Set technical direction for the data platform by identifying ETL bottlenecks and proposing pipeline/warehouse architecture improvements adopted across the team.
  • Mentored and ramped 10 offshore analysts on internal data tooling and workflows, cutting ramp-up time 40% and improving data-QA collaboration.
  • Partnered cross-functionally with product managers and analysts to deliver validated, analytics-ready datasets, reducing ad-hoc data requests and improving reporting consistency org-wide.

Tech stack

  • Python
  • SQL
  • Airflow
  • dbt
  • AWS S3
  • Glue
  • Redshift
  • Step Functions
  • n8n
  • OpenAI API
  • Gemini

Software Quality Analyst

BTI Solutions

  • Tested and validated hardware-software integration for in-vehicle telematics modules used by Hyundai MOBIS, ensuring reliability of connectivity and embedded system functionality.
  • Performed root-cause analysis on device and firmware failures, reproducing issues with Python and Node.js debugging tools and collaborating with engineering to resolve defects.

Associate Software Developer

Tech Mahindra

  • Maintained and enhanced backend services for British Telecom's Product Pipeline Reporting platform, sustaining 99% system uptime.
  • Developed backend features and fixes using Node.js and SQL, translating client requirements into production-ready solutions.

Skills

Languages
  • Python
  • SQL
  • Scala
  • JavaScript
Orchestration & Transformation
  • Airflow
  • dbt (models, tests, documentation)
  • n8n
Cloud & Data Warehousing
  • AWS S3
  • EC2
  • Glue
  • Redshift
  • Athena
  • Step Functions
  • Hive
Streaming
  • Apache Kafka
  • Spark Streaming
AI / LLM Data Engineering
  • OpenAI API
  • Gemini LLM
  • Prompt-driven extraction & classification pipelines
  • TensorFlow
  • scikit-learn
Databases
  • MySQL
  • PostgreSQL
  • MongoDB
Data Modeling & Quality
  • Dimensional / analytics-ready schema design via dbt
  • dbt data tests
Tools
  • Git
  • JIRA
  • Tableau
Data Engineering Concepts
  • ETL / ELT
  • Data Pipelines
  • Data Warehousing
  • Data Modeling
  • Dimensional Modeling
  • Batch Processing
  • Real-Time Processing
  • Data Quality
  • Distributed Systems
  • Cloud Computing

Projects

Real-Time Data Streaming Pipeline

A distributed streaming pipeline that ingests and processes web-server logs in real time for low-latency traffic analytics, with scalable event ingestion deployed on AWS and Elasticsearch for query and monitoring.

  • Apache Kafka
  • Spark Streaming
  • AWS Kinesis
  • S3
  • Elasticsearch

Newsfeed Classifier

A newsfeed aggregator and classifier pipeline triggered every 30 minutes. It extracts content from financial news sources and news APIs, then applies a custom machine learning classifier to label each story as Funding, M&A, IPO, or Noise. Hosted on AWS EC2, it emails daily updates through AWS SES and archives labeled data in S3.

  • Python
  • Machine Learning
  • AWS EC2
  • SES
  • S3
View on GitHub

Next Word Predictor

A Django web app built on GPT-2's text generation model. It reads the context of the sentence being typed and offers three predicted next words, each with a probability score.

  • Python
  • Django
  • GPT-2
View on GitHub

EDGAR Web Scraping

A Python script that pulls mutual fund holdings from SEC EDGAR for a given ticker or CIK, parses the HTML and XML filings, and writes the holdings out as a .tsv file.

  • Python
  • Requests
  • BeautifulSoup
View on GitHub

Education

Campbellsville University

MS, Computer Science

New Jersey Institute of Technology

MS, Electrical Engineering (Computer Systems Architecture)

Vellore Institute of Technology

BS, Electronics & Instrumentation Engineering

Get in touch

The fastest way to reach me is email.

venkatesh.mymail@gmail.com

+1 (437) 838-1965 · Toronto, ON