Sanchit Goyal

Backend & Distributed Systems Engineer · Tech Lead at PayPay

I build payment systems where being wrong is expensive.

Gurgaon, India / 10+ years experience / sanchitgoyal2604@gmail.com / LinkedIn / Résumé (PDF)

What I do

I'm a Tech Lead at PayPay, Japan's largest mobile payment service — 75 million registered users, listed on Nasdaq since March 2026 — where I own the multi-tenant communications platform behind PayPay, PayPay Card, PayPay Bank and PayPay Securities.

Ten years in, almost all of my work has been on systems that move money — wallets, point balances, refunds, reconciliation. That shapes how I build. A payments platform that is fast but occasionally inconsistent is worse than useless, so I care more than most about migration safety, idempotency, tenant isolation, and understanding how a system behaves under load before customers discover it for me.

I also set technical standards and run architecture reviews for 7–8 engineers, and I tend to pick up the debugging nobody volunteers for — a Java 21 virtual-thread issue took me a week to isolate and ship a fix for.

Selected work

Moving 25 TB off TiDB, live

PayPay · 2025

Our core wallet data sat on TiDB, and the licence and storage costs were growing faster than the business needed them to. I was the sole application engineer on the migration to AWS Aurora, from proof of concept through production release, working alongside two DBAs.

About 25 TB in total: 75M wallet rows, 225M account rows, and roughly 100 billion rows across four ever-growing transaction and idempotency tables. Dual-write replication kept Aurora current with TiDB, so the actual cutover needed a single 30-minute maintenance window rather than a copy of that size. Aurora caps tables at 64 TB, so I added live partition rotation and archival to contain future growth against that ceiling rather than discovering the limit later.

  • 25 TB, ~100B rows
  • 30-minute cutover
  • $1M annualized saving
  • no data-consistency incidents

A multi-tenant notification platform at 20,000/second

PayPay · 2025–2026

Four regulated financial entities — PayPay, PayPay Card, PayPay Bank and PayPay Securities — needed to send across push, email, SMS, in-app, LINE, WhatsApp and merchant webhooks without competing for the same capacity. I owned the platform architecture and the capacity work, and wrote the onboarding guidelines each sister company follows.

One constraint was isolation: one tenant's traffic spike must not degrade another's delivery. The other was the clock. Japan restricts promotional push outside daytime hours, and a same-day campaign has to land before noon — so the day's entire campaign volume has to clear in a few hours, and transactional sends can never queue behind it. That is why a platform averaging 35 million notifications a day is built for 20,000 per second: a 20M-recipient campaign inside half an hour needs about 11,000 per second on its own.

Kafka and queues, worker autoscaling, batching, rate limiting and per-tenant quotas took us from 4,000 to 20,000 notifications per second at 99.99% uptime. I also built end-to-end observability on Prometheus and Grafana that ties a single notification back to its campaign or transaction, so campaigns get CTR and CVR while transactional sends are held to delivery within 5 seconds of payment.

  • 35M notifications/day
  • 20,000/second peak
  • 99.99% uptime
  • 4 tenants, isolated

An LLM support agent, to 100% of users

PayPay · 2026

We evolved a rule-based chatbot into an LLM agent through proof of concept, pilot and gradual rollout to the entire user base. The architecture uses separate models for embeddings, classification and response generation. I own the RAG design, the knowledge-update pipeline and the agent-memory system.

In a financial product the hard part isn't fluency, it's restraint. Answers are grounded strictly in approved support knowledge, and genuinely high-risk situations — an unrecognised payment, say — are routed to a human agent rather than answered. Current scope is question answering; backend action execution is still being built.

  • P90 TTFT 5s
  • under ¥5 (~US$0.03) per conversation
  • 100% rollout

Shipping points that expire — and reconcile

PayPay · 2024–2025

PayPay Points (Limited-Time) are points that carry an expiry and are restricted to particular services. I owned the point-system implementation and the reconciliation for them.

Granting them is the easy half. The hard half is refunds and reconciliation: when a purchase made partly with expiring, service-restricted points is refunded, the system has to decide what to give back, in what form, against a balance that may since have expired — and the ledger has to still agree with itself afterwards. That's where most of the design effort went. The release earned a Rockstar Award.

  • 350 TPS peak grants
  • Rockstar Award

Untangling point balances from the wallet domain

PayPay · 2023–2024

PayPay Point accounts lived inside a complex wallet domain that had accumulated responsibilities. Working from an architect-defined high-level design, a seven-engineer team — including me — owned the low-level design, cross-team coordination, implementation and the migration itself.

My share was data and API migration, reconciliation, testing and production readiness. The payoff beyond throughput: point flows can now scale independently, and the wallet domain got meaningfully less tangled.

  • 750 → 1,000 TPS capacity
  • payment P99 150ms → 100ms
  • end-to-end P99 1.2s → 0.9s
  • no data-consistency incidents

Replacing the platform behind a US insurer's group benefits

Guardian Life · 2020–2023

Group Benefit Platform Transformation was a multi-year programme to replace the operating platform behind all of Guardian's group benefits products and services. The binding constraint was that the legacy systems had to keep serving clients throughout — there was no window in which claims could simply stop. I led and mentored 5+ engineers decomposing those platforms into microservices and decoupling them in phases, including the migration onto a new claims platform.

The integration layer was the interesting part: cloud services and on-premise systems had to stay coherent while ownership of a given capability moved from one to the other. I designed that architecture. Separately I built an Angular self-service configuration application on AWS Lambda, API Gateway and DynamoDB, so clients could configure their own setup rather than routing every change through support — worth $12K per client annually. Integrating a Drools rules engine then cut client unloading time in half.

  • $12K saved per client, annually
  • 50% faster client unloading
  • 5+ engineers led

Where I've worked

Apr 2024 – now

Tech LeadPayPay India — Japan's largest mobile payment service, Nasdaq-listed

Oct 2023 – Mar 2024

Senior Backend EngineerPayPay India

Aug 2021 – Sep 2023

Senior Lead EngineerGuardian India Operations — Guardian Life, US Fortune 250 insurance

Jan 2020 – Jul 2021

Lead EngineerGuardian India Operations

Aug 2019 – Dec 2019

Software EngineerOrange Business Services — Orange S.A.

Aug 2018 – Jul 2019

Business Systems AnalystOptum — UnitedHealth Group

Jan 2018 – Jul 2018

Software EngineerOptum

Jul 2016 – Dec 2017

Associate Software EngineerOptum

Tools

Languages & frameworks
Java, Kotlin, Python, Spring Boot
Data & messaging
Kafka, Amazon Aurora, MySQL, Cassandra, Redis, TiDB
Infrastructure
AWS, Kubernetes, Docker, Terraform, Prometheus, Grafana
Design
Distributed systems, microservices, multi-tenant platforms, live data migration
AI
LLM applications, RAG, agent design, agent memory

Education

B.Tech, Electronics & Communications Engineering — Greater Noida Institute of Technology (Dr. A.P.J. Abdul Kalam Technical University), 2016.