Venkatesh Sridharan
Data Engineer · Toronto, ON
Canadian Permanent Resident – no sponsorship required
I'm a data engineer with 7+ years building and owning production data platforms for
financial-intelligence and analytics products. I work hands-on across the AWS data stack
(S3, Glue, Redshift, Athena, Step Functions), Airflow orchestration, and dbt transformation
modeling, and I design LLM-orchestrated data pipelines – entity extraction,
classification, and enrichment workflows built on the OpenAI and Gemini APIs.
My track record is taking ETL systems from bottleneck to reliable at scale, and growing
team capability through hands-on mentoring.
- 7+ years in production data engineering
- 30% lower pipeline processing latency
- 25% better financial-event tagging accuracy
- 100+ analyst hours saved every week
Experience
Oct 2019 – Aug 2026
New York, NY (Remote)
-
Architected and owned ETL data pipelines and dbt transformation models
across AWS (S3, Glue, Step Functions, Redshift), cutting processing latency 30% and
raising the reliability of datasets feeding ML training and downstream analytics.
-
Designed and led development of an AI-driven company-profile ingestion
pipeline (n8n, Gemini LLM, Python) that transforms vendor datasets into structured
company profiles – descriptions, NAICS/PICS classification, keyword metadata –
automating onboarding of 10k+ companies per batch and materially cutting manual analyst
effort. Read the case study
-
Built and deployed a real-time news classification system ingesting
1K+ articles/day from 20+ sources (Python, Airflow, AWS), auto-tagging IPO, M&A,
Funding, and Delisting events and eliminating 100+ analyst hours/week.
-
Built LLM-powered entity extraction and enrichment pipelines (OpenAI
API, n8n), improving financial-event tagging accuracy 25% within 6 months and expanding
automated coverage of company-level financial signals.
-
Set technical direction for the data platform by identifying ETL
bottlenecks and proposing pipeline/warehouse architecture improvements adopted across
the team.
-
Mentored and ramped 10 offshore analysts on internal data tooling and
workflows, cutting ramp-up time 40% and improving data-QA collaboration.
-
Partnered cross-functionally with product managers and analysts to
deliver validated, analytics-ready datasets, reducing ad-hoc data requests and improving
reporting consistency org-wide.
Tech stack
- Python
- SQL
- Airflow
- dbt
- AWS S3
- Glue
- Redshift
- Step Functions
- n8n
- OpenAI API
- Gemini
Software Quality Analyst
BTI Solutions
Mar 2019 – Jul 2019
Troy, MI
-
Tested and validated hardware-software integration for in-vehicle telematics modules
used by Hyundai MOBIS, ensuring reliability of connectivity and embedded system
functionality.
-
Performed root-cause analysis on device and firmware failures, reproducing issues with
Python and Node.js debugging tools and collaborating with engineering to resolve defects.
Associate Software Developer
Tech Mahindra
Jun 2015 – Jun 2016
Bangalore, India
-
Maintained and enhanced backend services for British Telecom's Product Pipeline
Reporting platform, sustaining 99% system uptime.
-
Developed backend features and fixes using Node.js and SQL, translating client
requirements into production-ready solutions.
Skills
- Languages
-
- Python
- SQL
- Scala
- JavaScript
- Orchestration & Transformation
-
- Airflow
- dbt (models, tests, documentation)
- n8n
- Cloud & Data Warehousing
-
- AWS S3
- EC2
- Glue
- Redshift
- Athena
- Step Functions
- Hive
- Streaming
-
- Apache Kafka
- Spark Streaming
- AI / LLM Data Engineering
-
- OpenAI API
- Gemini LLM
- Prompt-driven extraction & classification pipelines
- TensorFlow
- scikit-learn
- Databases
-
- Data Modeling & Quality
-
- Dimensional / analytics-ready schema design via dbt
- dbt data tests
- Tools
-
- Data Engineering Concepts
-
- ETL / ELT
- Data Pipelines
- Data Warehousing
- Data Modeling
- Dimensional Modeling
- Batch Processing
- Real-Time Processing
- Data Quality
- Distributed Systems
- Cloud Computing
Projects
Case study · PrivCo
AI Company-Profile Enrichment Workflow
I led and built an n8n workflow that turns raw vendor company records into structured,
publishable company profiles. Each record arrives with a name, headquarters, and
several overlapping descriptions from different sources; detailed prompts I wrote
guide the LLM through three jobs per company.
-
Classify the company into PICS (PrivCo industry keywords) and NAICS.
-
Describe it with a single sanitized, verified company description.
-
Tag it with keywords that power search and group similar companies
together for comparison and competitor analysis, which is the core of PrivCo's
product.
Design decisions
-
One step per task, not one large prompt. Classification,
description writing and keyword tagging produce different outputs and fail in
different ways. Separate steps keep each prompt short and focused, and each one can
be tuned and retried without discarding good output from the others.
-
A closed taxonomy. The model chooses from the fixed PICS list,
supplied as a retrieval (RAG) document, and the NAICS list. Anything outside those
lists is rejected.
-
Structured output. Every step returns JSON, which is validated
before it moves on.
Keeping it accurate
-
Grounding rule. The prompts restrict the model to facts present in
the input record.
-
Description rules. No marketing language, no person names, no
first-person “we” phrasing, and the company name must be present.
-
Rule-based routing. Python rule checks run on the output. Records
that pass every check are auto-accepted; the rest go to analysts with the reason
attached.
-
Thin inputs are not guessed at. Records with empty or too little
source text are skipped and handled manually later.
Running at scale
-
Batching and retries. Companies are processed in batches of 10k,
sometimes 40–50k, with n8n retrying failed calls.
-
Resumable runs. Output is written as the run progresses, so an API
limit or an unexpected failure loses nothing. On restart, input IDs are compared
against the IDs already processed and the run continues where it stopped.
-
A separate processed-ID list. Re-opening the output files to find
finished companies was slow in n8n, so I keep the processed IDs in their own small
file, which makes the restart check fast.
-
Failure logs. Detailed, timestamped log files are written to AWS
S3, recording which records failed and why.
Results
- 2–3 days per batch including analyst verification, where 10k companies took about a month by hand
- 88% of classifications accepted by analysts without edits
- 40–50k companies in the largest batches
What I'd build next
Analyst verification is now the slowest part of a batch, and routing today relies on
basic rules. Stronger automated checks would send fewer records to analysts.
-
Grounding check. Extract the numbers, places and names from each
description and flag any that do not appear in the source record.
-
Classification confidence. Classify each company twice and
cross-check PICS against NAICS; disagreement marks the record as uncertain.
-
Golden set. Keep a few hundred analyst-approved companies to test
every prompt change against before it runs on a real batch.
- n8n
- Gemini LLM
- Python
- Prompt design
- RAG
- AWS S3
Case study · PrivCo
AI QA Pipeline
I built an automated data-quality pipeline to take routine checking off QA analysts.
Given a list of companies and the values PrivCo already holds for them, an AI research
agent checks each company on the web, compares what it finds with the stored record,
and decides which updates are safe to apply and which need a person.
-
Resolve the company's website: normalize the URL, try likely
variants, and use the first one that responds.
-
Research the company with an AI agent that searches the web and
reads pages for firmographics, financials and corporate events, returning strict
JSON with a confidence rating on every field.
-
Compare the findings with the existing record after parsing them
into typed values, so revenue figures, headcount ranges and years line up.
-
Triage every finding into one of five buckets: auto-apply, no
action, verify before update, human review, or analyst-only.
-
Report the results as JSON for the import process and as one
self-contained HTML report per company for analysts.
Design decisions
-
Confidence decides the route. Only a high-confidence value for a
field PrivCo has no data on is applied automatically. Anything that conflicts with
existing data goes to a person, with the reason attached.
-
Tolerance, not exact matching. Numbers within 5% of the stored
value count as a match, so small differences between sources do not create review
work.
-
One contract, two model providers. The research step sits behind a
single function interface, with one implementation on the Claude Agent SDK and one
on Gemini using search grounding and function calling. Every other module is
shared.
-
A controlled field vocabulary. Only known fields can be imported.
Anything else, including corporate events such as a missed acquisition, is flagged
for an analyst and never written automatically.
Cost and failure handling
-
Hard budgets per company. The Gemini path stops at 8 tool calls or
$0.06 per company, and the Claude path at 180 seconds.
-
Failing safe. If a site is unreachable or the agent times out, the
record is marked as needing a human; nothing is guessed.
-
Built to hand over. I documented the output contract, the triage
rules and how to extend each stage, so another engineer could pick the pipeline up.
- Python
- Claude Agent SDK
- Gemini API
- Search grounding
- Function calling
- httpx
Real-Time Data Streaming Pipeline
A distributed streaming pipeline that ingests and processes web-server logs in real time
for low-latency traffic analytics, with scalable event ingestion deployed on AWS and
Elasticsearch for query and monitoring.
- Apache Kafka
- Spark Streaming
- AWS Kinesis
- S3
- Elasticsearch
Newsfeed Classifier
A newsfeed aggregator and classifier pipeline triggered every 30 minutes. It extracts
content from financial news sources and news APIs, then applies a custom machine
learning classifier to label each story as Funding, M&A, IPO, or Noise. Hosted on
AWS EC2, it emails daily updates through AWS SES and archives labeled data in S3.
- Python
- Machine Learning
- AWS EC2
- SES
- S3
View on GitHub
Next Word Predictor
A Django web app built on GPT-2's text generation model. It reads the context of the
sentence being typed and offers three predicted next words, each with a probability
score.
View on GitHub
EDGAR Web Scraping
A Python script that pulls mutual fund holdings from SEC EDGAR for a given ticker or
CIK, parses the HTML and XML filings, and writes the holdings out as a .tsv file.
- Python
- Requests
- BeautifulSoup
View on GitHub
Education
Campbellsville University
MS, Computer Science
Jul 2021 – Jun 2023
Louisville, USA
New Jersey Institute of Technology
MS, Electrical Engineering (Computer Systems Architecture)
Aug 2016 – May 2018
Newark, USA
Vellore Institute of Technology
BS, Electronics & Instrumentation Engineering
Aug 2011 – May 2015
Vellore, India