Role Overview
Twosixtechnologies is hiring a Data Collection Engineer. This is a full-time remote role, with the team based in Remote, USA. posted last week. Full responsibilities, required qualifications, and the apply link are listed in the description below.
Resume Keywords to Include
Make sure these keywords appear in your resume to improve ATS scoring
Job description
At Two Six Technologies, we build, deploy, and implement innovative products that solve the world’s most complex challenges today. Through unrivaled collaboration and unwavering trust, we push the boundaries of what’s possible to empower our team and support our customers in building a safer global future.
Two Six Technologies is seeking a highly skilled Data Collection Engineer to design, scale, and maintain our distributed web scraping and data extraction infrastructure. In this role, you will be responsible for building resilient data pipelines that harvest data from complex web ecosystems, ensuring strict data quality through automated validation, and managing containerized workloads at scale. If you thrive on reverse-engineering web applications, overcoming anti-bot barriers, and orchestrating distributed systems, we want you on our team.
Location: 100% Remote
What you will Do:
- Distributed Crawler Development: Design and deploy high-performance, distributed web scrapers using Python and Scrapy to extract massive datasets efficiently.
- Dynamic Content Extraction: Utilize Browser Scripting tools to navigate, interact with, and extract data from modern, dynamic, and JavaScript-heavy websites.
- Infrastructure & Container Orchestration: Deploy, scale, and manage scraping workloads on Kubernetes, ensuring optimal resource allocation and fault tolerance.
- Data Validation & Quality Assurance: Define strict JSON Schemas and leverage Pydantic to enforce data types, validate incoming payloads, and catch data drift early.
- Data Ingestion & Storage: Build and optimize search and storage pipelines using Elasticsearch, transforming raw web dumps into highly structured, searchable data.
- Pipeline Workflow Management: Architect robust pipeline workflows to manage the end-to-end data lifecycle—from discovery and extraction to validation and storage.
- Anti-Bot & Proxy Engineering: Manage complex proxy rotation, session handling, and browser fingerprinting to maintain high success rates against advanced anti-scraping systems.
What you will need (basic qualifications):
- Experience: 7+ years of professional software engineering experience, with a heavy focus on web scraping, data engineering, or distributed systems.
- Analytical Mindset: Excellent reverse-engineering skills, with the ability to dissect network traffic, unearth hidden APIs, and bypass complex web barriers.
- Reliability Focus: A strong commitment to data integrity, system monitoring, and building self-healing scraping systems.
- Core Language: Expert-level proficiency in Python.
- Scraping Frameworks: Deep experience with Scrapy and distributed scraping architectures (e.g., handling distributed queues, broad vs. deep crawling).
- Automation & Browser Scripting: Proven experience with browser automation tools (Playwright, Selenium, or Puppeteer).
- Data Serialization & Validation: Mastery of JSON, JSON Schema, and data validation using Pydantic.
- Search & Analytics Engines: Hands-on experience indexing, querying, and optimizing Elasticsearch clusters.
- Orchestration: Strong proficiency in managing and scaling applications within Kubernetes environments.
- Workflow Management: Experience building structured pipeline workflows to handle complex, multi-stage data extraction tasks.
- Education: Bachelor’s degree in Computer Science, Engineering
Nice if you Have:
- AI & Intelligent Extraction: Experience leveraging LLMs or Computer Vision for adaptive scraping, parsing unstructured HTML, or bypassing CAPTCHAs (AI in data collection).
- Cloud Infrastructure: Strong hands-on experience with AWS ecosystems (e.g., EKS, EC2, S3, RDS).
- Relational Databases: Proficiency in SQL for querying, schema design, and storing structured relational data.
- In-Memory Data Structures: Experience with Redis (specifically for caching, deduplication, or as a Scrapy distributed queue back-end).
- Event Streaming: Familiarity with Apache Kafka for real-time data streaming and decoupled pipeline architectures.
- Containerization: Strong foundation in Docker for local development and containerizing scraping microservices.
- DevOps: Experience with CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins) for automated testing and deployment of crawlers.
Clearance Requirement:
- Eligible to obtain a clearance
Two Six Technologies is committed to providing competitive and comprehensive compensation packages that reflect the value we place on our employees and their contributions. We believe in rewarding skills, experience, and performance. Our offerings include but are not limited to, medical, dental, and vision insurance, life and disability insurance, retirement benefits, paid leave, tuition assistance and professional development.
The projected salary range listed for this position is annualized. This is a general guideline and not a guarantee of salary. Salary is one component of our total compensation package and the specific salary offered is determined by various factors, including, but not limited to education, experience, knowledge, skills, geographic location, as well as contract specific affordability and organizational requirements.
Looking for other great opportunities? Check out Two Six Technologies Opportunities for all our Company’s current openings!
Ready to make the first move towards growing your career? If so, check out the Two Six Technologies Candidate Journey! This will give you step-by-step directions on applying, what to expect during the application process, information about our rich benefits and perks along with our most frequently asked questions. If you are undecided and would like to learn more about us and how we are contributing to essential missions, check out our Two Six Technologies News page! We share information about the tech world around us and how we are making an impact! Still have questions, no worries! You can reach us at Contact Two Six Technologies. We are happy to connect and cover the information needed to assist you in reaching your next career milestone.
Two Six Technologies is an Equal Opportunity Employer and does not discriminate in employment opportunities or practices based on race (including traits historically associated with race, such as hair texture, hair type and protective hair styles (e.g., braids, twists, locs and twists)), color, religion, national origin, sex (including pregnancy, childbirth or related medical conditions and lactation), sexual orientation, gender identity or expression, age (40 and over), marital status, disability, genetic information, and protected veteran status or any other characteristic protected by applicable federal, state, or local law. For more information review the Two Six Technologies Equal Employment Opportunity and Affirmative Action Policy and the EEO Poster.
If you are an individual with a disability and would like to request reasonable workplace accommodation for any part of our employment process, please send an email to accommodations@twosixtech.com. Information provided will be kept confidential and used only to the extent required to provide needed reasonable accommodations.
Additionally, please be advised that this business uses E-Verify in its hiring practices.
By submitting the following application, I hereby certify that to the best of my knowledge, the information provided is true and accurate.
About Twosixtechnologies
Twosixtechnologies
60 other open roles at Twosixtechnologies on TryApplyNow.
Frequently Asked Questions
How do I apply for the Data Collection Engineer position at Twosixtechnologies?
Use the Apply button above to submit your application directly to Twosixtechnologies. Most applications take less than 5 minutes if your resume and contact details are ready, and you'll be routed to the employer's official application system to finish.
Is the Data Collection Engineer role at Twosixtechnologies remote?
Yes. This is a remote role. The team is based in Remote, USA, but the position itself does not require relocating to that office.
What does a Data Collection Engineer at Twosixtechnologies earn?
Twosixtechnologies has not disclosed a salary range in this posting. Many employers share specifics later in the interview process; you can also ask during a recruiter screen if compensation transparency is important to you.
When was the Data Collection Engineer role at Twosixtechnologies posted?
This role was posted on July 13, 2026 (10 days ago). It's still listed as actively hiring; we re-confirm openings against the source system multiple times per day and remove closed roles.
More Jobs at Twosixtechnologies
View all →Lead Vulnerability Researcher
Twosixtechnologies
Lead Software Reverse Engineer
Twosixtechnologies
Software Developer
Twosixtechnologies
Lead Cryptographer
Twosixtechnologies
Senior Technical Operations Engineer
Twosixtechnologies
AI-powered job search
Get every job scored to your resume
Upload your resume and get jobs ranked, your resume tailored, and employee contacts found automatically.
Get started freeNo credit card to start