Open to full-time roles · Noida, India

I build production systems, and the pipelines that judge the models that write code.

Software engineer building production systems and evaluating the models that write code. 2+ years on Node, Next.js and AWS. Currently benchmarking coding LLMs at Quess Corp.

Resume
years of experience
2+
startups
6
CGPA, B.Tech CSE
9.4
public repos
60+
Prince Raj portrait

Now

Benchmarking coding models against real PRs at Quess Corp

GitHub

01 · About

Evidence over opinion. Boring infrastructure.

Software engineer with 2+ years of experience shipping full-stack products on Node, Next.js and AWS, and the last year spent on LLM evaluation: designing rubrics, benchmarking coding models and building the internal platforms that score them. B.Tech CSE at Galgotias University (9.4 CGPA), graduating 2026.

How I work

Evidence over opinion

Every judgement gets a rationale. Scores without reasons are noise.

Boring infrastructure

ALB + ASG + a pipeline that deploys on push. Clever infra is a liability at 2 am.

Types at the boundary

Validate inputs once, trust them everywhere. TypeScript end to end.

Ship, then harden

Get the fuzzy version in front of people, then make it reliable and fast.

Measure before optimising

Profile first, then fix the slowest thing. Latency 1.86s to 1.2s came from numbers, not guesses.

Own the outcome

From empty repo to production and post-launch support. If it is live, it is mine to keep healthy.

Now

  • Benchmarking coding models against real PRs at Quess Corp
  • Going deeper on Kubernetes and observability
  • Strengthening system design and distributed systems fundamentals

Facts

  • B.Tech, Computer Science and Engineering, Galgotias University · CGPA 9.4 / 10 · Sep 2026
  • Noida, India · Remote, hybrid or on-site (NCR / Bengaluru). Open to relocation.
  • Open to: Full-Stack Engineer, Backend Engineer, AI / LLM Evaluation Engineer, Frontend Engineer

02 · Experience

Shipped at every stage, from first commit to production.

Full-time roles across product startups, from frontend architecture to AI evaluation infrastructure.

Dec 2022Present
2023202420252026

06 / 06 · LLM evaluation for coding models · client: Ethara AI

Quess Corp

Software Engineer Gurugram, India

Feb 2026 → Present
9 mocurrent

Evaluating frontier coding models against real engineering work: real GitHub PRs, golden patches, binary rubrics and side-by-side judgments that feed post-training.

  • Designed and ran structured evaluation workflows for coding LLMs by analysing real GitHub PRs, extracting task metadata (language, category, difficulty, critical files) and validating solutions against golden reference patches.
  • Wrote 5 to 10 task-specific binary rubrics per task (40 to 50% correctness-weighted) to grade outputs across correctness, code quality, behavioural reasoning and summary accuracy.
  • Ran side-by-side (SxS) evaluations of STEM and coding responses on a weighted framework (instruction following, truthfulness, correctness, writing style, verbosity) to determine model preference.
  • Benchmarked internal model outputs against GPT-5.2 and Gemini-3 Pro, produced evidence-based justifications and delivered structured JSON reports used in post-training analysis.
Feb 2026 – Present

03 · Work

Things that actually ship.

For each project: the problem, what I built and what changed.

SafeReport screenshot

Product

SafeReport

AI-powered anonymous crime reporting platform

Problem
Reporting an incident is slow, intimidating and rarely anonymous. People give up before the report is filed.
Outcome
A complete report in under a minute, tracked until it is resolved, with the reporter's identity never stored.
What I built
  • Gemini-powered incident analysis that classifies the report, extracts key facts and drafts the structured report automatically.
  • Real-time geolocation with Google Maps so the exact location is captured without typing an address.
  • End-to-end encryption for sensitive fields, JWT authentication and rate limiting to protect reporters and the platform.
  • Prisma schema on NeonDB tuned with indexes and query patterns to handle 10,000+ concurrent report submissions.
Next.jsTypeScriptPrismaNeonDB (Postgres)Gemini AINextAuthDockerTailwind
Live Code

Case study · internal at WhatBytes

LLM Evaluation Platform

The system a team used to decide which model is actually better

Problem
Comparing LLMs by vibes does not scale. The team needed repeatable, auditable scores across static datasets, live API runs and multi-turn conversations.
Outcome
Model decisions moved from opinion to evidence: every score has a rationale, every run is reproducible.
Next.jsTypeScriptNode.jsPostgreSQLRedisOpenAI / Gemini APIs
Internal project

Open source

Production-grade Backend

Typed Node/Express API with a real CI/CD path to AWS

Problem
Most starter backends stop at `npm run dev`. This one ships.
Outcome
Push to main, live on EC2 a few minutes later. Used as the base for my own services.
Node.jsExpressTypeScriptMongoDBDockerAWS ECR / EC2GitHub Actions
Code Template

Open source · most-starred

AWS Deployment with CI/CD Guide

From SSH + PM2 + Nginx to Docker + ECR + GitHub Actions

Problem
Deployment tutorials either stop at 'it works on EC2' or assume you already run Kubernetes.
Outcome
A guide other developers actually use; my most-starred repository.
AWS EC2NginxPM2DockerAWS ECRGitHub Actions
Code

04 · Try it

Grade a PR the way I do.

At Quess Corp I score coding-model outputs with task-specific binary rubrics, weighted toward correctness, then compare side by side. Here is a tiny real-style example. Grade it, then see my scoring and why.

TaskReturn the last N orders for a user, newest first. Default N to 10.
// orders.ts
−export async function recentOrders(userId: string, n: number) {
− const orders = await db.order.findMany({ where: { userId } });
− return orders.slice(0, n);
+export async function recentOrders(userId: string, n = 10) {
+ const orders = await db.order.findMany({
+ where: { userId },
+ orderBy: { createdAt: "asc" },
+ take: n,
+ });
+ return orders;
}

Model's own summary of the change

“Moved sorting and limiting into the database query, added a default limit of 10 and newest-first ordering.”

Binary rubric · 100 pts

50% correctness-weighted
  • Correctness · 30 pts

    Orders come back newest first.

  • Correctness · 20 pts

    Limiting happens in the database, not in memory.

  • Code quality · 15 pts

    Default parameter and query shape are idiomatic Prisma.

  • Behavioural reasoning · 15 pts

    Handles edge input (n <= 0, very large n) sensibly.

  • Summary accuracy · 20 pts

    The model's summary matches what the diff actually does.

05 · Skills

The toolkit, grouped by where it is used.

Production experience across the stack, with a current focus on evaluation infrastructure and platform reliability.

Programming Languages

TypeScriptJavaScriptPythonSQLC++

Application Engineering

ReactNext.jsReact NativeAngularTailwind CSSFramer Motion

Backend Systems & Data

Node.jsExpress.jsPostgreSQLMongoDBPrismaRedisGraphQLREST APIs

Cloud Infrastructure & Delivery

AWS (EC2, S3, CloudFront, Lambda)DockerKubernetesGitHub ActionsCI/CD pipelinesServerless

LLM Evaluation & Applied AI

Rubric designSxS evaluationModel benchmarkingMulti-turn scoringLLM APIsPrompt engineering

Architecture & Reliability

MicroservicesEvent-driven systemsDistributed systems designObservabilityZero-downtime deployments

06 · People

What people I've worked with actually wrote.

From founders, leads and engineers I have worked with. Each one links to LinkedIn.

I had the pleasure of working with him on multiple projects, where he showcased a strong understanding of React.js and Next.js concepts while delivering high-quality code. He consistently demonstrated a willingness to learn and grow, and his enthusiasm for tackling new challenges was truly inspiring. Highly recommended!
Abhijit Agarwal

Abhijit Agarwal

Founder, Xipper

I had the privilege of working alongside Prince on multiple industrial projects, and his talent and work ethic truly stood out. His professionalism and attention to detail were remarkable. I wholeheartedly endorse Prince for any opportunity that requires a highly skilled and dedicated individual.
Sathya Sachi Paira

Sathya Sachi Paira

CEO, SecWebXperts

Had the pleasure of working with Prince during his internship at Xipper. Quick to learn, made essential contribution with developing the MVP, and showed strong initiative throughout. A reliable and thoughtful developer anyday!
Shyan Roy Choudhury

Shyan Roy Choudhury

SDE-2, Publiq Studio

A talented React frontend developer with an exceptional eye for responsiveness. His proficiency in Node.js further enhances his capabilities, enabling him to build robust and scalable applications. Attention to detail, problem-solving skills and commitment to delivering outstanding results.
Ayush Kumar Tiwari

Ayush Kumar Tiwari

Software Engineer, D4 Community

He did a great job during his time with us, especially in building clean and user-friendly web screens for our project at Torus. Prince is hardworking, quick to learn, and always delivered his tasks on time.
Deepanshu Kamboj

Deepanshu Kamboj

Full Stack Developer, Devtown

He's quick to understand requirements, writes clean and efficient code, and always brings a positive attitude to the team. Whether it's tackling tricky UI bugs or building smooth user interfaces, Prince handles it all with focus and care.
Harshit Bafna

Harshit Bafna

Full Stack Developer, Akal Infotech

07 · Proof

Public track record.

GitHub activity, competitive programming profiles and certifications.

github.com/HardCoder404

public repos

…

stars earned

…

followers

…

Problem solving

1800+

LeetCode problems solved · 3-star

1200+

Codeforces rating

200+

day LeetCode streak

GSSoC'24

open-source contributor

AWS Academy Graduate · Cloud Foundations
AWS Academy Graduate · Cloud Architecting
C++ · Certificate of completion

08 · Contact

Hiring, or just curious?

I reply to every real message, usually within a day. Email is fastest.

+91 82183 28333

Open to full-time roles

Full-Stack Engineer · Backend Engineer · AI / LLM Evaluation Engineer · Frontend Engineer

Remote, hybrid or on-site (NCR / Bengaluru). Open to relocation.

No spam, no newsletter. Just a reply.