PRACTICE GUIDE TWO SIGMA
Two Sigma Coding and Data Analysis Practice Test
The Two Sigma coding and data analysis screen is 3 questions in 180 minutes - about 60 minutes each - written in a code editor and run against tests you cannot see. There is no negative marking, so leaving an item blank gains you nothing over guessing.
Where it sits: Coding and data analysis round. The practice sitting on this page runs the same item count, the same clock and the same marking rule, with questions generated by Outcry rather than taken from Two Sigma.
Outcry is not affiliated with Two Sigma and has no access to their assessment content. This guide describes an assessment format that candidates report publicly; the questions here are generated by Outcry and are not Two Sigma’s own.
What it screens
A quantitative investment manager applying machine learning and distributed computing to markets.
- ✓Reading concurrent code and pointing at the defective line
- ✓Order-book data structures: what happens on a cross, a cancel, a partial fill
- ✓Algorithmic complexity chosen for the input sizes stated in the problem
- ✓Monte Carlo methods and when a simulation beats a closed form
Where it sits at Two Sigma
This is the whole sitting rather than one section of it: 3 items in 180 minutes, on one clock.
You can return to earlier items within the section, so a first pass for the quick ones and a second for the rest is a workable plan.
The format
These are the numbers the Two Sigma sitting on this site runs on, matching the format candidates report.
| Questions | 3 |
|---|---|
| Time | 180 minutes |
| Per question | 60 minutes |
| Negative marking | No |
| Answer style | Code editor, run against hidden tests |
| Where it sits | Coding and data analysis round |
What it tests, with a worked example
Every example below is generated by Outcry, drawn from the same question generators the timed drills run. None of them is Two Sigma’s.
Data structure implementation
Build the thing rather than call it: an order book, a cache, a ring buffer. Graded by hidden tests, so there is no partial credit for an approach.
Complexity that is actually graded
Hidden tests sized so the naive solution times out. A correct answer that is too slow scores the same as a wrong one.
Reading someone else's code
A diff or a function with a bug in it, and the question is where.
Example
A hot loop does `results.push_back(x)` a million times with no other setup. What is the cheapest fix?
- reserve() the final size once before the loop
- Switch to std::list to avoid copying
- Make results a static global
- Use emplace_back instead of push_back
Answer reserve() the final size once before the loop
Growth reallocates and copies roughly log₂(n) times. One reserve makes it a single allocation. emplace_back saves a construction, not the reallocations.
Language and systems detail
Memory, references, undefined behaviour and the things that bite in production.
Example
A counting semaphore starts at 5. 7 completed wait operations and 5 completed signal operations have run against it. What is its value now?
- 2
- 3
- 10
- -2
Answer 3
Each completed wait subtracts one and each signal adds one: 5 - 7 + 5 = 3. Only completed waits count, which is why a negative reading would mean blocked processes rather than a legal value.
Numerical and data handling
Floating point, aggregation and joins on data that does not fit the obvious shape.
Example
A fund has volatility 20% and the index has volatility 24%. Their correlation is 0.8. What is the fund's beta to the index?
Answer 0.6666666666666666
Beta = correlation x (fund vol / index vol) = 0.8 x 20/24 = 0.667. Correlation and beta only agree when the two volatilities match, which here they do not.
Reading a spec precisely
Return type, ordering and tie-breaking are graded, and are where most silent failures come from.
What a good score looks like
On a paper of 3 questions with no penalty for a wrong answer, the only thing an unanswered question can do is cost you. Coding screens are usually pass-fail on hidden tests rather than scored, and candidates commonly report that a solution passing every correctness test still fails on a timeout. Treat full marks as solving every problem inside the complexity bound, not merely solving it.
How to train for it
- 01Implement the core structures from scratch once each. The screens ask you to build them, not use them.
- 02Write the edge cases into your own tests first. That is where the hidden ones live.
- 03Get the function signature exactly right before anything else. Return type and ordering are graded, and a correct algorithm behind a wrong signature scores zero.
TRAIN IT HERE
The drills that match each section
Concurrency Clash
Find the data race, name the fix - real C++ defect patterns, five levels.
Order Book
Matching-engine mechanics: crosses, cancels and partial fills.
Algorithm Lab
DP tables, Monte Carlo estimation and speed rounds at four difficulty tiers.
SIT THE FULL BATTERY
All the sections back to back on one clock, marked the way the real screen marks them, with a by-skill breakdown at the end. Included with any pass.
Also reported at Two Sigma
Common questions
- Is the Two Sigma coding and data analysis test multiple choice?
- Code editor, run against hidden tests. You write and run code against tests you cannot see.
- How long is the Two Sigma coding and data analysis test?
- 3 questions in 180 minutes, which is about 60 minutes each.
- Is there negative marking on the Two Sigma coding and data analysis test?
- No. A wrong answer costs nothing beyond the mark you would have earned, so leaving an item blank is never better than guessing at it.
- How do I practise for it free?
- Every drill linked on this page is free to play, with no account, inside a daily run cap. Questions are generated fresh each run, so there is nothing to memorise between attempts. The full Two Sigma Quant Research Online Assessment sitting puts the sections back to back on one clock.