LexCore: Intelligent Document Intelligence Platform
Scenario
Sarah Chen, CTO — LexCore Legal Technologies
“Our legal team is losing 3–4 hours every day to manual document search. We have 500,000 internal documents — case files, contracts, compliance guides, and legal precedents — stored in S3. I need our engineers to build a system where a lawyer types a question in plain English and gets a precise, cited answer drawn from our actual documents in under 2 seconds. Not from the internet, not hallucinated — from our documents only. We have three clearance tiers: Junior Associate, Senior Counsel, and Partner. A junior associate must never be able to retrieve Partner-level documents, even if those documents are technically relevant to their query. At peak, all 500 of our lawyers may be querying simultaneously. This must be production-grade, fully observable, and fault-tolerant — not a proof of concept.”
Requirements
- 1All queries must enter through a secure API gateway. No user may reach the document store or vector index directly.
- 2The document source is classified and must never be publicly accessible or reachable by end users.
- 3Access control must be enforced at retrieval time, filtering results by clearance tier: Junior Associate, Senior Counsel, or Partner.
- 4The LLM must be grounded. Responses must cite retrieved document chunks only and never draw from training data.
- 5Documents must be parsed, chunked (≤ 512 tokens, ≤ 40 token overlap), and embedded before storage in the vector index.
- 6System must serve 5,000 concurrent queries at peak with p95 response time ≤ 2 seconds.
- 7Query latency, retrieval quality, and LLM response accuracy must be instrumented and observable in production.
Skills Assessed
Challenge Details
5,000
Concurrent users
500
Requests / second
10 min
Time limit