Back to Jobs
Consultant - Data Architect
Actively Reviewing
HCA Healthcare - India
Job Description
Position Summary
The Staff Data Architect is a hands-on technical role responsible for designing and personally building the data architecture patterns that power HCA’s GCP-based Lakehouse and multi-modal ingestion platforms. Unlike a purely advisory architecture role, this position is expected to spend a substantial portion of its time in code — authoring reference implementations, prototyping ingestion frameworks, writing production-grade pipelines, and proving out patterns before they are adopted at scale.
This role owns architecture for a defined platform or data domain rather than the enterprise as a whole, and operates as the senior technical authority within that scope. The Staff Data Architect translates business and clinical requirements into durable designs, then demonstrates those designs in working code that engineering teams extend. Design decisions are earned through implementation: if a pattern cannot be built, benchmarked, and operated by the architect who proposed it, it is not ready to become a standard.
The role bridges the Principal Data Architect’s enterprise direction and the Staff Data Engineer’s delivery execution. It requires deep GCP expertise, strong software engineering discipline (test-driven development, CI/CD, infrastructure-as-code), and the ability to influence without direct authority. The culture of the organization places an emphasis on teamwork, so social and interpersonal skills are equally important as technical capability. Due to the emerging and fast-evolving nature of GCP and AI technologies, the position requires staying well-informed of technological advancements and being proficient at putting new innovations into effective practice.
At HCA DT&I, your deliverables will influence patient care. Every process, technology, and decision matters.
Major Responsibilities:
Core Competencies
The following are highlighted entrepreneurial competencies and core expectations for the job/role:
Bachelor's degree in computer science, related technical field, or equivalent experience
Required
Master's degree in computer science or related field
Preferred
7+ years of experience in Data Engineering, Data Architecture, or equivalent hands-on data platform development
Required
2+ years of experience in a Cloud Data or Information Architect capacity
Required
3+ years of experience in Healthcare
Preferred
8+ years of experience in Information Technology
Required
Knowledge, Skills, Abilities, Behaviors:
A Successful Candidate Will Demonstrate:
○BigQuery, Cloud Storage, Dataflow, Dataproc, Cloud Composer
○Pub/Sub, Kafka, Spark Streaming
○Cloud Run, GKE, Cloud Functions
○Lakehouse formats: Iceberg, Delta Lake, Parquet, Avro, JSON
○CI/CD pipelines, Git/GitHub, test-driven development
○Bigtable, Cloud SQL, Cloud Spanner
○Microservices architecture, OpenAPI/REST and event-driven API design
○Vertex AI, LangChain, embeddings, RAG patterns
○Healthcare domain knowledge/experience
The Staff Data Architect is a hands-on technical role responsible for designing and personally building the data architecture patterns that power HCA’s GCP-based Lakehouse and multi-modal ingestion platforms. Unlike a purely advisory architecture role, this position is expected to spend a substantial portion of its time in code — authoring reference implementations, prototyping ingestion frameworks, writing production-grade pipelines, and proving out patterns before they are adopted at scale.
This role owns architecture for a defined platform or data domain rather than the enterprise as a whole, and operates as the senior technical authority within that scope. The Staff Data Architect translates business and clinical requirements into durable designs, then demonstrates those designs in working code that engineering teams extend. Design decisions are earned through implementation: if a pattern cannot be built, benchmarked, and operated by the architect who proposed it, it is not ready to become a standard.
The role bridges the Principal Data Architect’s enterprise direction and the Staff Data Engineer’s delivery execution. It requires deep GCP expertise, strong software engineering discipline (test-driven development, CI/CD, infrastructure-as-code), and the ability to influence without direct authority. The culture of the organization places an emphasis on teamwork, so social and interpersonal skills are equally important as technical capability. Due to the emerging and fast-evolving nature of GCP and AI technologies, the position requires staying well-informed of technological advancements and being proficient at putting new innovations into effective practice.
At HCA DT&I, your deliverables will influence patient care. Every process, technology, and decision matters.
Major Responsibilities:
Core Competencies
The following are highlighted entrepreneurial competencies and core expectations for the job/role:
- Architectural judgment grounded in hands-on implementation
- Deep technical expertise in cloud data platforms
- Strong problem-solving, systems design, and software engineering skills
- Ability to bridge business, clinical, and technical domains
- Communication and interpersonal skills; influence without authority
- Understand strategic imperatives
- Design data architecture patterns for a defined platform or domain within the GCP Lakehouse ecosystem, and personally build the reference implementations that prove them.
- Write production-quality code (Python, SQL, Java/Scala) for ingestion frameworks, transformation pipelines, and platform tooling — not prototypes thrown over the wall, but modular, tested, reusable code that serves as the pattern others follow.
- Benchmark and validate architectural options empirically: build competing approaches, measure cost, throughput, and latency, and let evidence drive the standard.
- Spend meaningful time in the codebase alongside engineers — pairing, reviewing pull requests, and debugging production issues — to keep designs anchored in operational reality.
- Produce architecture decision records (ADRs), reference architectures, and diagrams that are backed by working artifacts in source control.
- Modernize and refactor legacy pipelines, leading migrations through direct contribution rather than delegation alone.
- Design and build Lakehouse environments using BigQuery, Apache Iceberg, Delta Lake, and GCS, including schema evolution, ACID compliance, partitioning, clustering, metadata management, and lifecycle policies.
- Model data and access patterns that serve both analytical and AI-driven workloads; implement the initial models and validate them against real query workloads.
- Optimize storage and compute for cost and performance across large-scale document and data repositories; instrument pipelines so cost and performance are observable, not estimated.
- Define and uphold SLAs for timely delivery of data, and build the monitoring that proves they are met.
- Design, build, and tune ETL/ELT and real-time ingestion pipelines using Dataflow, Dataproc, Pub/Sub, Kafka, Spark Streaming, Cloud Run, Cloud Composer, GKE, and Cloud Functions.
- Build multi-modal ingestion capabilities handling structured, semi-structured, and unstructured data (documents, text, images, PDFs, audio/video metadata).
- Implement AI-assisted ingestion patterns — semantic chunking, embedding generation, metadata extraction, classification, and enrichment — in partnership with the AI/ML practice, and integrate them into pipelines hands-on.
- Partner with AI/ML teams on vector storage and retrieval (RAG) patterns aligned with enterprise governance standards.
- Ensure ingestion frameworks are resilient, observable, idempotent, and designed for continuous evolution.
- Design and implement REST and event-driven APIs aligned with OpenAPI specifications, versioning standards, and backward compatibility principles.
- Apply enterprise API governance — naming conventions, authentication patterns, rate limiting, pagination, and error handling — and contribute the shared libraries and templates that make compliance the default.
- Build API layers serving both human-facing applications and AI agent consumers, ensuring consistent contract design across downstream integrations.
- Implement secure API access patterns for PHI-sensitive and regulated data, incorporating encryption, scoping, and audit logging.
- Review and validate API designs from vendor and internal delivery teams prior to implementation.
- Champion test-driven development, continuous integration, and automated deployment by building them into the platforms this role delivers.
- Implement unit and integration tests; conduct performance testing where appropriate.
- Build and maintain CI/CD pipelines and infrastructure-as-code (GitHub, Terraform) for data platforms.
- Implement automated workflows that lower manual/operational costs and move the company closer to democratizing data.
- Enable self-service data architecture supporting query exploration, dashboards, data catalog, and rich data discovery.
- Apply and enforce best practices for data governance, security, privacy, and compliance (HIPAA, GDPR) across structured and unstructured data.
- Ensure designs align with enterprise policies for data retention, lineage, access control, and auditability — and implement the controls, not just the policy.
- Participate in and lead architectural design reviews to ensure adherence to standards and patterns.
- Collaborate with business, clinical, analytics, and engineering stakeholders to translate requirements into scalable solutions.
- Partner with the Principal Data Architect to align domain architecture with enterprise direction, and provide implementation feedback that shapes enterprise standards.
- Mentor Staff and Senior Data Engineers through code review, pairing, and design coaching; raise the technical ceiling of the team by example rather than by directive.
- Lead data analysis efforts and solution proposals to data-related and data architecture problems.
- Be a leader in the HCA data community. Evangelize architecture and engineering best practices, participate or present at community events, and encourage the continual growth and development of others.
- Be curious. Be growth minded. Encourage and enable this in others.
- Demonstrate professional and personal maturity through self-leadership.
- Build productive and healthy relationships within the department and other teams to foster growth of our culture, our people, and our platforms.
- Practices and adheres to the “Code of Conduct” philosophy and “Mission and Value Statement.”
- Perform other duties as assigned.
Bachelor's degree in computer science, related technical field, or equivalent experience
Required
Master's degree in computer science or related field
Preferred
7+ years of experience in Data Engineering, Data Architecture, or equivalent hands-on data platform development
Required
2+ years of experience in a Cloud Data or Information Architect capacity
Required
3+ years of experience in Healthcare
Preferred
8+ years of experience in Information Technology
Required
Knowledge, Skills, Abilities, Behaviors:
A Successful Candidate Will Demonstrate:
- Hands-on experience designing and building enterprise data solutions on Google Cloud Platform or another major cloud provider — with a portfolio of systems personally built, not only specified.
- Must-have skills/tools:
○BigQuery, Cloud Storage, Dataflow, Dataproc, Cloud Composer
○Pub/Sub, Kafka, Spark Streaming
○Cloud Run, GKE, Cloud Functions
○Lakehouse formats: Iceberg, Delta Lake, Parquet, Avro, JSON
○CI/CD pipelines, Git/GitHub, test-driven development
- Nice-to-have skills/tools:
○Bigtable, Cloud SQL, Cloud Spanner
○Microservices architecture, OpenAPI/REST and event-driven API design
○Vertex AI, LangChain, embeddings, RAG patterns
○Healthcare domain knowledge/experience
- Experience with document and unstructured data processing, including ingestion, enrichment, and indexing.
- Practical experience integrating LLMs and AI frameworks into production data pipelines.
- Strong understanding of data security, privacy, and regulatory requirements (HIPAA, GDPR) in cloud environments.
- Strong ability to assemble large, complex data sets meeting functional and non-functional requirements.
- Strong ability to identify, design, and implement internal process improvements, including redesigning data platforms for greater scalability and optimized data delivery.
- Expert ability using source control management and CI/CD automation tools.
- Strong understanding of Agile methodologies and how to apply Agile within the team.
- Ability to communicate complex architectures clearly to both technical and non-technical audiences, and to present and facilitate technical ideas.
- Proven ability to complete work, make sound decisions, and plan and accomplish goals without explicit direction/guidance from leadership.
- Demonstrates an empathetic and growth mindset with a willingness to learn new skills, technologies, and methodologies.
- Coaches and mentors engineers within and external to the team.
- Excellent problem-solving and analytical skills.
- GCP Professional Data Engineer
- GCP Professional Cloud Architect
Required Skills
Similar Jobs
View all →
Sr. Data Analyst
RSCP
Noida
₹6 LPA–₹9 LPA
Dashboard Development
Data Analysis
ETL Validation
+1
SAP ABAP Development for HANA
Aryvart Software Pvt Ltd
Coimbatore
₹5 LPA–₹17 LPA
SAP ABAP on HANA
HANA
SAP
+1
PHP Developer
Cloudnausor Technologies Pvt ltd
Chennai
₹4 LPA–₹10 LPA
PHP
SQL
Senior Backend Engineer
Ctruh
Bengaluru
Event-driven architecture
Data architecture
OAuth 2.0
+36
Senior Data Engineer (GCP)
Pixeldust Technologies
Mumbai
Retrieval-Augmented Generation
Machine Learning
Adobe Illustrator
+18
Share
Quick Apply
Upload your resume to apply for this position
–