A complete example
Everything CaseBrick turns a single posting into — the tailored questions, one answer coached out loud, the delivery read, and a readiness number. One example, every step, nothing cut.
One posting in. Optionally your résumé, so questions target the gap between what the role wants and what you've done.
Senior Deep Learning Software Engineer, TensorRT
NVIDIA · AI Inference (TensorRT / TensorRT-LLM) · Santa Clara, CA or US Remote · onsite loop
Build and optimize the software that runs the world's largest models in production. Day to day, you will:
What we need to see: 4+ yrs performance software · expert C++/Python · a DL framework (PyTorch, JAX, or TensorFlow) · GPU architecture & CUDA (or Triton, CUTLASS) · Linux, Docker.
Ways to stand out: LLM-inference optimization (quantization, KV-cache, paged attention, speculative decoding) · compiler/MLIR background · OSS contributions.
Base range $184,000–$339,250 USD · equity · full benefits
CaseBrick reads what the posting actually screens for and writes the set — behavioral and role-specific, targeted to your gaps, not a generic bank.
Generated from common public patterns for this role and level — not real or leaked questions, and CaseBrick tells you so.
“I've spent four years getting models into production — the last two obsessing over inference latency, not just accuracy. I haven't shipped a custom TensorRT kernel, but I've profiled and cut real GPU serving costs, and I go deep fast.”
You answer out loud. CaseBrick scores what you said against a structured-interview rubric and how you sounded on six delivery signals — then hands you the single fix that moves the needle most.
Um, first thing — I profile before I touch anything. I'd run Nsight Systems to see whether I'm compute-bound or memory-bandwidth-bound, because the fix is completely different. For LLM decode you're usually bandwidth-bound moving the KV cache, so I'd batch more requests to amortize the weight loads, then, like, reach for FP8 quantization to shrink what I'm actually moving. If attention is the bottleneck I'd swap the naive kernel for a fused, FlashAttention-style one. And I'd hold every change to a fixed latency budget and watch quality on a held-out set, so a speedup doesn't quietly regress accuracy.
Two hedges ("um," "like") early cost you steadiness — costly in a systems interview, where they read your first sentence for conviction. Open on the profiler result, not "um, first thing." A decisive opener reads as senior.
Every rehearsed answer updates one readiness number for this role — it climbs as you sharpen, and it can dip if you add a harder round or leave an answer weak. It's honest, not a vanity meter.
Interview-ready
Started at 44 after one rep → 79 after six. One systems answer still flagged to sharpen.
A practice estimate — it doesn't predict or guarantee your interview result.
How the scores are built and what they can't tell you: Delivery Engine · Signal Layer.
Paste the posting you're actually preparing for and get your own set back in seconds.
Try freeNo signup to see your questions · audio scored on-device, never sent to us