Search results
0 results. Even S3 finds things faster than that — try “terraform”, “security”, “observability”, or “lambda”.
Real numbers, pulled from CloudWatch & DynamoDB when you loaded this page — not a screenshot.
Projects
build → ship → operate. Click any tile to open its full write-up (abstract, war stories, links).
More on GitHub — C++ systems & concurrency experiments at github.com/Abheenash.
Skills
The stuff on the résumé — minus the buzzword bingo.
Experience & Education
Where the XP came from. Pick a section, then a company.
Application support for the workplace systems a refinery runs on day to day — the applications that decide whether a contractor gets through the gate, whether a work permit is valid, and whether a facility request gets actioned. The estate spans AWS and on-premises RHEL, written over the years in Java, C++, Perl and Ruby.
- Support 7 workplace applications used by approximately 13,000 employees and contractors across refineries, terminals, and offices, covering contractor access, work permits, and facility requests. The applications run on both AWS and on-premises RHEL, so the same week can involve a CloudWatch alarm on one side and a cron job on a RHEL box on the other.
- Work approximately 20 production tickets a month. Each one starts with evidence — Splunk logs and SQL against Oracle and SQL Server — to establish what actually happened before anything is changed, so fixes address the cause rather than the symptom and the ticket doesn't come back.
- Fixed 12 production defects across Java, C++, Perl, and Ruby. The most involved was a C++ memory leak in the site-access service that had been managed with weekly restarts for long enough that the restart had become routine; I reproduced it, isolated it with Valgrind and GDB, and shipped a peer-reviewed patch, so the service no longer needs the restart.
- Built the monitoring that shortened detection. The on-premises Perl and C++ estate had no alerting of its own, so I added Splunk alerts and dashboards for it; for the AWS workloads I added CloudWatch alarms managed in Terraform, so the alarm definitions live in code and are reviewed like any other change; and I introduced structured error logging so failures are searchable rather than buried in free text. Together they caught 3 failures before any user reported them.
- Partner with the application developers on enhancements rather than only fixes, including bulk work-permit renewals in a legacy Ruby on Rails application, so permits can be renewed in batches instead of one at a time.
- Hold the on-call rotation one week in five. During an outage I run the recovery procedures and coordinate across development, plant IT, and the service desk so each group knows what is happening and what it owns; between incidents I maintain the runbooks and present incident trends at the monthly service reviews.
- Use GitHub Copilot and Amazon Q for log triage and script drafting — a useful first pass over a large log or a first draft of a script — with every generated suggestion reviewed before it ships.
- Contractor records live in three systems — Oracle, ISNetworld, and SAP — and when they drift out of sync a contractor who should be cleared gets stopped at the gate. The workaround was a manual pre-shift check, done by hand before every shift, to catch mismatches before people arrived.
- Built a nightly Java reconciliation job covering approximately 3,500 active contractor records across all three systems. It repairs the mismatches it can resolve automatically, attaches SQL diagnostics to the ones it can't, and escalates those through ServiceNow, so a person only ever looks at the genuinely ambiguous cases. It runs on RHEL cron under Control-M scheduling.
- The manual pre-shift check was retired, and team-wide gate-access tickets fell from approximately 12 a month to 4 — a number that held through refinery turnarounds, when contractor volume is at its highest.
A summer on the systems side of a large internal-services estate: who owns what, whether alerts actually reach the right people, and what to do when the answer is “nobody”.
- Built a Go service auditing on-call ownership and alert routing across approximately 120 internal services, exporting the gaps as Prometheus metrics so they could be graphed and alerted on like any other signal. It surfaced 31 services with missing or stale owners — services where an alert would have fired to no one, or to someone who had moved on.
- Delivered the audit with a runbook for resolving each kind of ownership gap and a Grafana dashboard that tracks the gaps over time, so the number keeps going down after the internship rather than drifting back up.
- Investigated service-catalog discrepancies with read-only SQL against PostgreSQL and ClickHouse and found approximately 1,800 orphaned ownership records. Wrote a cleanup procedure for them with a tested rollback, so the fix could be reversed if it turned out to be wrong.
- Shadowed the on-call rotation for 6 weeks, triaging Prometheus and Grafana alerts alongside a mentor, and coauthored 2 incident reports — the first time I'd written up an incident for an audience beyond my own team.
- Raised test coverage on an internal Go service above 70% and added CI checks so it stays there, and containerized 2 Python tools for Kubernetes so they run the same way everywhere instead of on one person's machine.
Two years of cloud operations for a U.S. client's B2B platform on AWS — on-call, incidents, and the automation that made both quieter.
- Supported a U.S. client's B2B platform on AWS serving approximately 300 business customers and 3 million API requests daily, across development, staging, and production. Troubleshot compute, load balancing, database, IAM, VPC, and OS issues across ECS Fargate, RDS, and Linux.
- Held the on-call rotation one week in four: triaged CloudWatch and PagerDuty alerts, assessed customer impact, executed documented rollbacks and recovery procedures, and authored more than 10 root-cause analyses so the same incident didn't recur.
- The alarms were static thresholds, which page on noise. Replaced them with golden-signal alarms — latency, traffic, errors, saturation — and composite service-health alarms, each tied to a runbook, cutting pages per on-call week from approximately 30 to 11.
- Automated the recurring support work with Python, Boto3, Lambda, EventBridge, and Systems Manager — patch-compliance reporting, non-production scheduling, and resource health checks — saving approximately 4 hours a week, and remediated more than 60 AWS Config, GuardDuty, and Security Hub findings.
- Releases were half manual — some steps automated, the rest done by hand in the console — which made each release slow and each production change hard to trace afterwards.
- Rebuilt the process into an automated GitHub Actions pipeline with tests, security scanning, image versioning, staged deployments, a gated production rollout, and automatic rollback if the rollout fails.
- Release time fell from approximately 2 hours to under 20 minutes, and console-based production changes ended: every production change now goes through the pipeline and leaves a trail.
- Approximately 150 AWS resources had been provisioned by hand, so environments drifted apart and standing up a new one took days of clicking.
- Moved them into reusable Terraform modules with remote state, state locking, pull-request plans — so every change is reviewed as a plan before it is applied — and drift detection.
- Configuration drift was eliminated, and new-environment setup dropped from approximately 3 days to 90 minutes.
- Relevant coursework: Principles of Internetworking · Introduction to Cybersecurity · Open Systems · Advanced Computer Architecture
Certifications
The paper trail. All ✓ verified — no “in progress” fillers here.




Self-paced, employer-authored job simulations via Forage — completed as hands-on skills practice (not employment).
Leave feedback
Spotted a bug, have a role in mind, or just want to say the console theme is unhinged? It lands straight in my inbox — via a Lambda I built, naturally.