Sr Manager IC FinOps

Date PostedAugust 13, 2026LocationRemote - EligibleCompanyCapital OneSalaryRemote (Regardless of Location): $209,000 - $238,500 for Sr. Lead Software Engineer; McLean, VA: $229,900 - $262,400 for Sr. Lead Software EngineerTypeSenior Level

Job Summary

Own the cost and AI-leverage layer of the analytics platform behind Capital One Shopping as an individual-contributor role. Build net-new systems including a cost-attribution pipeline, a forecast-driven autoscaler control loop, and a production Gen AI system. Lead engineering efforts with deep technical expertise in data platforms and cloud infrastructure.

Responsibilities

  • Own and evolve data warehousing, event ingestion, and Airflow orchestration production platforms.
  • Build a cost-attribution pipeline parsing Trino query logs.
  • Develop a forecast-driven autoscaling control loop for shared Trino clusters and dbt worker pools.
  • Design and ship a production Gen AI system with RAG and natural-language-to-SQL capabilities.

Required Skills

  • 10+ years of engineering experience in production environments
  • Proficiency in Kafka, streaming technologies, and large data volumes
  • Expertise in SQL, Python, Java, Go, TypeScript/JavaScript, and AWS ecosystem
  • Production Gen AI and RAG design experience
  • Claude Code fluency

Job Details

The role Own the cost and AI-leverage layer of the analytics platform behind Capital One Shopping - the systems between petabyte-scale data and the humans and tools that query it. This is an own-and-build role, not a maintenance seat: you own the production platforms below, and in your first six months you ship three net-new systems on top of them. You'll report to the Engineering Director for the Shopping data platform as one of two senior IC pillars of the analytics org. What you own You own, support, and evolve three production platforms - ideation through implementation to production support - and you're the SME and mentor for the analysts, BAs, and engineers who use them: - A data warehousing platform serving ~250K queries/day over ~20PB. - An event ingestion pipeline taking in 6-7 billion events/day. - The Airflow orchestration platform. You own the technology choices and the strategic backlog and priorities for this surface - and you carry ongoing production support and an on-call rotation for these platforms and the models on them. You build the three net-new systems below on top of that operational base. What you'll build - A cost-attribution pipeline that parses Trino query logs at production traffic and attributes real AWS dollars to every report, query, user, and dbt model - reconciled against a ~$750K/month cloud bill. - A forecast-driven autoscaling control loop for the shared Trino cluster and dbt worker pool - turning today's event-only Nomad autoscaler (Prime Day, Cyber Week) into steady-state, forecast-driven capacity. The single largest lever on the analytics AWS bill. - A production Gen AI system - natural-language-to-SQL or RAG over the data catalog - with real LLM tool-use, grounding, and cost guardrails, adopted by internal teams. The stack Kafka streaming backbone into an S3 lakehouse (Hive + Iceberg), Cassandra, Postgres, DynamoDB, ElasticSearch, Aurora MySQL. Queried through Trino/Presto and Spark SQL, modeled in dbt, orchestrated on Airflow and Nomad and containers (Docker/Kubernetes), on a deep AWS footprint. SQL and Python daily; Go, Java, and TypeScript/JavaScript across the surrounding platform. The day-to-day "Manager" is the level, not the job - this is an individual-contributor role, and you'll spend most of your day hands-on in the editor. Roughly 70% building: writing the log-parsing and cost-attribution logic and its dbt models, building and tuning the forecast-driven autoscaler control loop, and building the RAG / natural-language-to-SQL system yourself. The other ~30% is technical coordination - reconciling your cost numbers with Finance, the R&D memo, aligning report owners - not status decks or people-management. Daily rhythm is multi-terminal Claude Code: query-log analysis in one, dbt work in another, AI iteration in a third. No direct reports - you build. What we're looking for - 10+ years engineering experience owning and supporting mission-critical applications and platforms in production - Deep experience with Kafka and streaming technologies and platforms - Experience with enterprise data technologies and platforms - A track record working with extremely large traffic and data volumes - Fluency across a real stack: JavaScript, Java, HTML/CSS, TypeScript, SQL, Python, and Go, open-source RDBMS and NoSQL databases, container orchestration (Docker and Kubernetes), and a broad range of AWS tools and services - Forecast- or workload-driven infrastructure scaling on a shared platform - Production Gen AI - RAG design, vector/graph stores, LLM tool-use - 0-to-1 delivery of a product that booked measurable revenue or adoption in its first weeks - Experience directing a cross-team migration or deprecation at senior-executive scope to completion - Comfortable working with large teams on large-scale systems, and cross-functionally - Claude Code fluency - daily use, skill authoring, PR-level deliverables (hard requirement) Basic Qualifications: - Bachelor's Degree - At least 6 years of experience in software engineering - At least 1 year experience with cloud computing (AWS, Microsoft Azure, Google Cloud) Preferred Qualifications: - Master's Degree - 9+ years of experience in at least one of the following: JavaScript, Java, TypeScript, SQL, Python, or Go - 4+ years of experience with AWS, GCP, Microsoft Azure, or another cloud service - 4+ years of experience in open source frameworks - 1+ years of people management experience - 2+ years of experience in Agile practices