AS
I'm Anjana,I design Azure Databricks lakehouses, medallion data platforms, and AI/ML-ready foundations — 10+ years, enterprise-grade, governed by design.
With deep expertise in data infrastructure and a sharp eye for reliability, I specialize in building pipelines, lakehouses, and real-time platforms.
Led modernisation of Cloudera big data workloads onto Azure Databricks, Delta Lake, ADF, ADLS Gen2 and Synapse. Processing speed improved 45% and infrastructure cost reduced 35%.
View ProjectBuilt streaming and micro-batch ingestion patterns using Kafka/Event Hub concepts and PySpark Structured Streaming, delivering near real-time analytical use cases on the lakehouse.
View ProjectOptimised terabyte-scale PySpark and Spark SQL pipelines through partitioning, caching, broadcast joins and query tuning — delivering 40-50% performance improvements alongside a 35% infrastructure cost reduction.
View ProjectImplemented Unity Catalog governance with catalog, schema, table and column-level controls for sensitive and PII datasets, plus validation frameworks covering schema enforcement, null and duplicate checks, reconciliation and freshness.
View ProjectAre you ready to see more?
See All WorksMy working process revolves around an approach aimed at maximising productivity and delivering clean, reliable data from day one.
Learn MoreAudit existing data landscape, identify pain points, map stakeholder needs and define success metrics.
Blueprint the data model, pipeline topology, and storage strategy. Choose the right tools for scale and budget.
Engineer robust pipelines with built-in data quality checks, monitoring, and CI/CD from day one.
Tune performance, reduce cloud spend, and iterate with the team to ensure the platform meets high standards of quality.
These numbers reflect the experience and consistency, and measurable impact behind the work I've delivered over the years.
I am Anjana, an Enterprise Data Architect based in Dubai, UAE. With 10 years spanning Azure, Databricks and Cloudera, I thrive at the intersection of cloud data platforms, governance, and modern engineering practice.
My journey across Accenture and Lumen Technologies has been a testament to my passion for crafting reliable, enterprise-grade data platforms and fearlessly pushing the boundaries of modern data engineering.
The services I offer are meticulously crafted and tailored to cater specifically to your unique needs and requirements.
Module lead for enterprise Azure Databricks lakehouse platforms. Defined Medallion bronze/silver/gold standards across ADLS Gen2, Delta Lake, ADF and Synapse, and led Cloudera-to-Azure modernisation that cut processing time 45% and infrastructure cost 35%. Own governance via Unity Catalog, PySpark performance tuning, and production operations as Azure Platform Admin.
Designed PySpark ETL on Cloudera to transform operational data into centralised warehouse and analytical layers. Built RDD and DataFrame pipelines with broadcast joins, partition pruning and caching, reducing processing time by up to 50%, plus Hive/Impala validation frameworks and Apache Oozie workflows for reliable production scheduling.
Led design and implementation of Oracle and Informatica PowerCenter data warehousing solutions, consolidating multi-source data for analytics and reporting. Improved query performance by 40% through indexing, partitioning and SQL tuning, and maintained Linux shell automation for pipeline orchestration, log analysis and error recovery.
Designed and developed ETL solutions for pharmaceutical and healthcare accounts, aligned to HIPAA and FDA regulatory expectations. Implemented Informatica PowerCenter pipelines across EHR, billing and operational sources, and optimised complex SQL queries to cut execution time by 35%.
Got a data challenge in mind? Let's build something great. I can transform that problem into a reliable, scalable platform.