Back to all workNext Project: Early Career Foundation

Strengthening Proctor Mode Integrity

Catching what mattered — without treating every candidate like a suspect.

A coding assessment is only worth something if the result can be trusted. Proctor Mode is the layer that flags suspicious behaviour during a remote test — but an integrity system that cries wolf is as useless as one that misses cheating entirely. The real problem was never “catch more cheating.” It was giving a recruiter a signal they could act on without wrongly accusing an honest candidate.

I designed the integrity experience end to end: the candidate’s onboarding into a proctored test, the in-test warnings when something looked off, and the recruiter’s review of what actually happened — turning raw events into evidence a person could reason about.

Role
**Lead Product Designer**
Timeline
**2018–2020 (built in phases)**
Team
PM · **Engineering** · HR · **Sales · Marketing**
2
Sides of the assessment served — candidate & recruiter
3
Shipped surfaces — onboarding, warnings, recruiter review
Signal
Not a verdict — integrity as evidence, never an accusation

Screens are original design mockups from the project. Metrics are described qualitatively where source data isn’t independently verifiable today.

HackerRank proctor mode: candidate onboarding and the real-time monitoring panel.
The shipped integrity experience. Below: how it got there — and every decision behind it.

Why did Proctor Mode need strengthening?

But before I could strengthen anything, I had to understand what ‘integrity’ was quietly failing to deliver — and for whom.

A candidate sits down to a coding test in their bedroom. No proctor in the room, no invigilator pacing the aisle — just a browser window and the honor system. On the other side, a recruiter has to look at the result and decide whether to trust it enough to stake an interview, a shortlist, a hire.

That gap — between a test taken anywhere, by anyone, under any conditions, and a hiring decision that has to feel safe to make — is where this project lived.

In 2021, at the peak of remote hiring, I led the design of HackerRank's foundational proctoring experience as the sole designer: the candidate's side of being fairly monitored, and the recruiter's side of trusting what came back. The brief was never to build a surveillance machine. It was to make a remote assessment feel fair to the person taking it and trustworthy to the person reading it — at the same time.

Not another lockdown cage. An invisible layer of trust: strong enough to protect the result, light enough that an honest candidate barely feels it.

In-person assessment worked because of things nobody had to design. One person, one screen, one room, a proctor who could see the whole space. The integrity of the test was held by the physical world itself — quietly, for free.

When hiring went remote in 2020–2021, that scaffolding vanished overnight. The test moved into living rooms, onto unmanaged laptops, across time zones — and every assumption in-person proctoring had handled for free now had to be handled by software, or not at all.

HackerRank already had raw safeguards: full-screen, tab tracking, copy-paste logging. But raw safeguards aren't an experience. They left candidates anxious about being watched without knowing how, and left recruiters holding a pile of technical events with no way to read them. The signal existed. The meaning didn't. That gap was the brief.

The gap was clear once I saw it. What wasn’t clear yet was exactly what I’d been asked to protect. So I pinned down the mandate.

What was actually at stake here?

No tidy brief. A live integrity system, real retention stakes, and a scope I had to define before I could defend it.

This wasn't a side feature. Evaluation is where HackerRank's customers get the thing they pay for — a hiring call they can act on. If a remote result can't be trusted, the entire funnel above it is wasted, and the customer goes looking for a platform they can rely on. Trust in the assessment was, quietly, a retention problem.

So the bet was simple to state and hard to earn: make a remote assessment something a candidate trusts is fair, and a recruiter trusts is real — without either one paying for the other's confidence.

I framed it internally as a trust problem, not a detection problem. Detection catches the dishonest few. Trust is what you owe the honest many — and it's what keeps a customer on the platform. Design only for the first and you corrupt the experience for everyone; design for both and you protect the business. That dual mandate became the principle every later decision answered to.

Owning it meant deciding what fair, trustworthy integrity even looked like. So I set the rules first.

What rules did every decision have to pass?

I didn’t write these down at the time. But the same handful of principles kept deciding the hard calls.

I didn’t formalise these as a list at the time. They emerged from the reframe — once integrity became a trust problem rather than a detection one, the same handful of rules kept deciding the hard calls, and every screen in this case study traces back to one of them.

Screen from the Strengthening Proctor Mode Integrity case study.
Screen from the Strengthening Proctor Mode Integrity case study.
Screen from the Strengthening Proctor Mode Integrity case study.

The rules were clear. What I didn’t have yet was evidence of where integrity actually broke. So I went and looked.

How did you learn what both sides actually needed?

The first real work was listening — understanding real assessments before proposing any fix.

I didn't start with screens. I started with the two people on either side of a remote test, and the same broken moment seen from opposite ends.

Working with my product manager and customer-facing teams, I gathered the recurring complaints from recruiters who lived in these reports daily and the anxieties candidates carried into a monitored test. None of it was a formal study — it was the practical, in-the- room understanding you build by listening closely to the people closest to the workflow. It was enough to point clearly at what the experience had to fix.

What candidates feltWhat recruiters needed
Anxious about being watched without knowing how or whyConfidence the result was earned fairly
Afraid an innocent action — checking docs — would look like cheatingA way to read integrity signals without forensic work
No idea what the rules were until they'd broken oneEvidence clear enough to act on, scoped enough to stay fair

The thing that kept coming up. From the recruiter side, the same frustration surfaced again and again: a report full of security events with no way to tell a four-second glance away from a four-minute disappearance. From the candidate side, the mirror image — people flagged for something they hadn't known was against the rules until the flag appeared. Neither side needed more data. They needed the events to carry context. That single observation is what the in-test warning loop and the recruiter timeline were both built to answer.

Underneath all of it was a single tension: candidates needed transparency, recruiters needed certainty, and most proctoring tools bought the second by sacrificing the first. That tension became the thing I designed against.

What I heard didn’t just confirm the problem — it moved it. ‘Catch the cheat’ turned out to be the wrong question.

How did “catch the cheat” become the wrong question?

The research settled the problem. Then it quietly rewrote it.

The core move was a reframe. Proctoring is usually built as a detection problem — catch the cheat. I decided it had to be a trust problem — earn the honest candidate's confidence, and give the recruiter something they could believe. That single shift changed every downstream choice.

I stated the problem across four lenses, to keep a wide stakeholder set honest about whose problem this actually was:

LensThe problem, stated
WhomCandidates being remotely assessed, and the recruiters judging the result
WhatTrust a remote test is fair, and read its integrity confidently
WhereInside the assessment — onboarding, the test itself, and the report
WhySo honest candidates aren't punished by anxiety, and recruiters aren't left guessing

Which let me frame two years of decisions as one question:

How might we help a remote candidate prove their work is their own, and a recruiter trust that it is — without turning the test into surveillance?

That HMW was the tie-breaker. Anything that protected integrity but eroded the candidate's trust failed it; anything that calmed the candidate but left the recruiter guessing failed it too.

Once I knew what integrity really had to protect, the next question was who it protected — and who it put at risk. Two people, opposite fears.

Who were the two people this had to work for?

Two people, one integrity signal, opposite fears — and the system only worked if it answered both.

The experience served many stakeholders, but two roles drove every decision. They wanted opposite things from the same system, and it only worked if it satisfied both. These are composite pictures drawn from the research above, not individuals.

Screen from the Strengthening Proctor Mode Integrity case study.
Screen from the Strengthening Proctor Mode Integrity case study.

The candidate’s fear and the recruiter’s doubt were the same problem from two sides. A test the candidate found hostile would produce a worse signal; a signal the recruiter couldn’t read would waste the candidate’s honest effort. Carry both through the anxious middle, and the system works.

Knowing who I was designing for, I could trace where their trust would break. So I mapped the whole journey.

Where did candidates actually struggle most?

Before changing a single screen, I walked the whole assessment — hunting the exact moments trust was won or lost.

I mapped the candidate's path from invite to submission to find where trust was won or lost. The shape told me where to spend: the emotional low sits in the middle — granting intrusive permissions, then risking a flag mid-test — exactly where a candidate decides whether this feels fair.

StageMINDSETWhat's happeningDesign opportunity
InviteCurious, a little waryReceives the test linkA calm, low-anxiety entry point
OnboardExposed — "what are they watching?"Grants webcam, screen, monitor checksExplain every ask in plain language
In-testTense — "did I just break a rule?"Codes; may trip a flagWarn gently, allow recovery
SubmitRelievedFinishes; identity capturedClose the loop, no ambush

The celebratory ends were easy. The valley — onboarding and the in-test middle — is where most of the design went. Carry the candidate through that, and the recruiter side never starves for trustworthy results.

The map showed me where the risk lived. Now it was time to design past it. Structure first, screens next.

How did you get from structure to screen?

No heavy wireframe stage. The integrity logic had to be mapped before it could be drawn.

The most important structural decision was that integrity couldn't be a bolt-on. It had to thread through the test surface the candidate and recruiter already used — part of the assessment, not a separate destination.

Screen from the Strengthening Proctor Mode Integrity case study.

Two branches, one surface. The candidate experience and the recruiter experience are two views of the same assessment, sharing the same underlying events. The only node that isn't mine is marked: the AI verdict layer that now reads from the recruiter timeline came later, and it computes its verdicts from the very signals this foundation was built to capture.

A structure on paper is a hypothesis until you trace the paths through it — including the ones that don't go well. I designed the unhappy branches with the same care as the happy path, because that's where a trust experience is actually won or lost.

Screen from the Strengthening Proctor Mode Integrity case study.

Two flows mattered most. The permission flow had to handle refusal gracefully — explaining why an access was needed and offering a retry, rather than dead-ending a candidate on an unsupported browser. And the in-test warning loop had to let an honest candidate self-correct: flag, warn, acknowledge, continue — with the event recorded to the recruiter timeline as context, never as a silent strike. Designing that no-flag loop and that denied branch as carefully as the success path is the whole argument of this project in one diagram.

The map showed me where the risk lived. Now it was time to design past it. Structure first, screens next.

What did you actually ship?

Restrained on purpose. An integrity signal has to read as fair before it reads as anything else.

The foundation of a fair test is set before the first line of code. I designed onboarding as a sequence of clear, plain-language steps — each permission gated and explained, so consent was informed rather than blind.

I proved the structure in greybox before any visual existed. The test was simple: could a candidate understand what was being asked, and in what order, with nothing styled to lean on? If the sequence reads in grey, it reads.

The Design · Onboarding

The foundation of a fair test is set before the first line of code. I designed onboarding as a sequence of clear, plain-language steps — each permission gated and explained, so consent was informed rather than blind.

I proved the structure in greybox before any visual existed. The test was simple: could a candidate understand what was being asked, and in what order, with nothing styled to lean on? If the sequence reads in grey, it reads.

From that skeleton, the high-fidelity flow kept the same spine: four gated permissions, each explained, in a fixed order that ends at full- screen.

Screen from the Strengthening Proctor Mode Integrity case study.

The pre-test permissions flow — webcam, monitors, screen share, full-screen — each explained before the test begins.

Current HackerRank product, shown for reference — the flow and interaction model are mine from 2021–2022; the visual skin and proctor mascot are later work.

The interaction pattern under every step is the same: a step explains itself, then asks. The collapsed row states what's needed; expanding it gives the reason and the action — explain, then request, never request blind.

Screen from the Strengthening Proctor Mode Integrity case study.
Screen from the Strengthening Proctor Mode Integrity case study.
Screen from the Strengthening Proctor Mode Integrity case study.

I led this as the sole designer, but the boundaries of what the browser could honestly observe were set with engineering — a constraint that became a design principle, covered in The Hard Parts.

Warnings & Recovery

When a candidate does something the rules flag, the response should correct gently, not accuse — the difference between a proctor's quiet word and a security guard's hand on the shoulder.

Screen from the Strengthening Proctor Mode Integrity case study.

The tab-switch warning — name the behavior, explain why it matters, let the candidate continue.

Current HackerRank product, shown for reference — the flow and interaction model are mine from 2021–2022; the visual skin and proctor mascot are later work..

Screen from the Strengthening Proctor Mode Integrity case study.

One integrity moment sits before the test even starts: confirming the person is who they claim to be. Impersonation is the one cheat that voids everything downstream — so identity verification was foundational, and I designed it to be done with the

Screen from the Strengthening Proctor Mode Integrity case study.

Making Signals Legible

The hardest problem wasn't capturing events — it was making them mean something to a recruiter who isn't a forensic analyst. I designed the recruiter side around a single idea: integrity evidence belongs on a timeline, anchored to when it happened and

Screen from the Strengthening Proctor Mode Integrity case study.
Screen from the Strengthening Proctor Mode Integrity case study.

The recruiter timeline — events placed in time, with duration and context.

Current HackerRank report, still built on the timeline structure I designed. The signals shown — monitor changes, tab and window activity — are the foundational layer.

The work was real and in people’s hands. The harder question was whether it changed anything. So here’s what moved.

What’s changed since you built this?

Integrity work rarely ships as one launch moment — and its purpose was never a single number.

Since 2022, HackerRank has built substantially on this foundation — most visibly by adding an AI layer. I show it here plainly, not to claim it, but because the fact that the foundation absorbed a whole new generation of capability is the strongest thing I can say about the structure.

Screen from the Strengthening Proctor Mode Integrity case study.

Today's integrity settings. Secure Mode — full-screen, copy-paste, monitors, tab alerts — is the foundational layer; the AI Add-on badges mark the later capabilities.

Screen from the Strengthening Proctor Mode Integrity case study.

The current AI-era session replay — a richer descendant of the recruiter-timeline idea, with verdicts computed from the same signals the foundation captured. Current product reference; the AI session replay is later work.

The newer verdict vocabulary — None, Medium, High — and the model-driven detection (HackerRank reports the plagiarism model at 85% precision) sit on top of the same events the foundational layer was built to capture. The AI added a judgment layer to the architecture, rather than replacing it.

What was the hardest part?

The honest version of this includes the parts that didn’t come easy — and the edge cases that decided everything.

Every honest case study has a second half. This is mine — and it's also where I draw the line on what's mine and what isn't, because the boundary is part of the story.

My first visual direction was rejected. The onboarding treatment I originally designed didn't ship — it was turned down during that period for a different look. It taught me what actually endures in design work: not the pixels, the structure. The flow and interaction model I set are what the product still runs on, through more than one visual reskin since — including the current proctor mascot, which isn't mine. My original 2021 files weren't preserved, so the screens here are current-product reference and the diagrams are my reconstructions; I'd rather be plain about that than pass off the present-day UI as my untouched work.

The browser's blind spots were real. Browser-based proctoring can't see a second device or a person off-camera. I designed an experience honest about that ceiling rather than performing a coverage it didn't have — security theater would only have added anxiety without adding protection.

The integrity-versus-anxiety line never fully resolved. Every safeguard was a cost to an honest candidate. I held that anything which only added stress without protecting the result got cut — but it was a judgment call every time, not a formula. And I'm deliberate about outcomes: this phase predates the AI-era product, so I claim no metrics. The one figure that circulates — 85% plagiarism precision — belongs to the later AI layer, not to me. What I stand behind is the structure, and its durability is the proof.

A clean happy path is the easy part. A trust experience is only trustworthy if it holds up when things go wrong — and most of the design judgment lived in the states a polished mockup never shows.

The messy caseHow I designed for it
Permission deniedExplain why the access is needed and offer a retry — refusal is a moment to build trust, not punish.
The false positiveThe hardest case: an honest candidate flagged. Warnings name the behavior and allow recovery; events reach the recruiter as context to interpret, never as an automatic verdict.
Unsupported browser or dropped connectionDegrade gracefully and tell the candidate plainly what to do — a broken session must never read as a cheating signal.
The over-flagging trapIf every candidate is warned constantly, the signal becomes noise and the experience turns hostile. I kept the flag threshold meaningful so a warning still meant something.

The throughline: every failure state had to stay fair to the candidate and honest to the recruiter at once. That balance was the whole job — and it never resolved into a formula. It was a judgment call every time.

Naming the hard parts leads to the last, most personal question. What did building this teach me?

What did this teach you?

I went in strengthening a feature. I came out understanding what it means to design systems that judge people fairly.

This project taught me to treat trust as a design material — something you don't ship in a release but earn through a product that behaves fairly every time, especially in the moments a user fears most. Designing the warning and the permission ask with as much care as the success path is the clearest thing I carried out of this work.

It also taught me that the durable part of design is the structure, not the surface. My first direction was rejected and the skin has been redesigned more than once — yet the bones I set are still load-bearing years later, now carrying an entire AI layer I never built. That reframed what I think I'm actually paid to get right.

And it sharpened how I want to work: honest about what's verified and what isn't, honest about what the technology can and can't do, and willing to design for the honest majority even when it would be easier to optimize for catching the dishonest few.

That’s what it built in me. And it’s what I bring to every high-stakes problem now.

Next case study

Early Career Foundation

Where the craft was forged — twelve years at service companies.

Let's connect.

I'm open to senior product design roles at product-led companies. If you're hiring, or want to talk through a role — I'd love to hear from you.

raghavendrashet@me.com