SHIVA.TRIPATHI_

Software Engineer · Kathmandu, Nepal

I turn messy
web data
into datasets
you can trust.

5+ years extracting, cleaning, and shipping structured data from dynamic websites, APIs, and unruly HTML — plus the automation and testing chops to make sure it stays reliable in production.

pipeline.py — extraction tier
$source ~= find_source(domain)
$try: fetch(source.api)
except NoAPI: BeautifulSoup(source.html)
except Rendered: Selenium(source.dom)
$normalize(regex, xpath, fuzzywuzzy)
dataset.ready — 0 malformed rows

5+

years in production data pipelines

95%

automated API coverage, up from 20%

3

extraction tiers: API → BS4 → Selenium

4h→3m

runtime cut on a legacy SQL workflow

Toolbelt

What I reach for

Extraction

PythonSeleniumBeautifulSoupRequestsXPathRegex

Processing

PandasFuzzyWuzzyETLData NormalizationSnowflake

Quality & Delivery

API TestingPostmanAppiumBrowserStackAzure Pipelines

Selected work

Recent builds