Task: Find Duplicate Records in Large Dataset with Limited Memory
Task: Find Duplicate Records in Large Dataset with Limited Memory: a task in Terminal-Lego-15k (Harbor dataset). You are given a large dataset on disk containing data records (approximately 1 KB per record). Your task is to implement a solution in C++ to find all duplicate records while working…
The task
You are given a large dataset on disk containing data records (approximately 1 KB per record). Your task is to implement a solution in C++ to find all duplicate records while working under strict memory constraints.
Part of PrimeIntellect/Terminal-Lego-15k.