InterviewLM is our own software — a multi-tenant SaaS platform we build and operate. Below is the product itself: a full video walkthrough, screenshots of each core surface taken from the running application, the feature set, and exactly where we are in development.
Both links work without an account.
A guided tour through creating a pipeline, the candidate's live interview, and the evaluation report that comes out the other end.
Video not loading? Watch it on Loom.
Screenshots below are captured from the running application.

The candidate debugs real code with an AI pair partner they are told to use. Note what the AI does here: asked whether the fix belongs in get() or put(), it pushes the reasoning back rather than handing over the answer. Every message, edit and run is recorded as a session event, which is what makes the score auditable later. This particular screen is our public five-minute sample — you can take it yourself, no signup. The full coding assessment uses the same layout with a file tree and terminal added.

When a candidate finishes, the platform produces a recommendation with its reasoning, the conditions a hiring manager should probe next, and an overall score. This is the real report surface, rendered from our production components — read the whole thing on our sample report page.

Four scored dimensions, each carrying the evaluator's confidence and the number of logged session events supporting it. AI collaboration is a first-class dimension — the part legacy platforms cannot measure because they ban AI outright. A candidate can score 91 on collaborating with AI and 74 on code quality, and the report says so instead of averaging it away.
What the platform does, concretely.
Each candidate gets an isolated container with a real filesystem, package installs, and a PTY terminal — not a text box that diffs against expected output. Problems are repository-shaped: debug this failing service, refactor this module, review this PR.
Candidates are told to use the AI. We score how well: prompt specificity, whether they verify output, whether they override the AI when it is wrong, and whether they can still reason independently. Prompt-delegators and real engineers separate here.
A voice agent conducts structured technical and behavioural rounds with no scheduling and no interviewer time. Candidates in any timezone start the moment they are invited.
Every session is reconstructable: keystrokes, terminal commands, AI prompts, test runs, and audio, on one timeline. A hiring manager who disagrees with a score can watch what actually happened instead of arguing about it.
Build a loop from 13 round types — resume screen, voice screen, coding, system design, work session, case study and more — across 5 role families, then let candidates auto-advance on the rules you set.
Scores are backed by an append-only event store, every evaluation is logged for bias audit, and integrity ceilings cap scores when the evidence is too thin to justify them. The report tells you when it is unsure.
InterviewLM is live and running paid assessments for real hiring teams — not a prototype or a waitlist. Figures below are as of August 2026.