{"version":"0.1","company":{"name":"YubHub","url":"https://yubhub.co","jobsUrl":"https://yubhub.co/jobs/skill/web-crawler"},"x-facet":{"type":"skill","slug":"web-crawler","display":"Web Crawler","count":4},"x-feed-size-limit":100,"x-feed-sort":"enriched_at desc","x-feed-notice":"This feed contains at most 100 jobs (the most recently enriched). For the full corpus, use the paginated /stats/by-facet endpoint or /search.","x-generator":"yubhub-xml-generator","x-rights":"Free to redistribute with attribution: \"Data by YubHub (https://yubhub.co)\"","x-schema":"Each entry in `jobs` follows https://schema.org/JobPosting. YubHub-native raw fields carry `x-` prefix.","jobs":[{"@context":"https://schema.org","@type":"JobPosting","identifier":{"@type":"PropertyValue","name":"YubHub","value":"job_9be280f4-cbc"},"title":"Software Engineer, Data Infrastructure","description":"<p>We&#39;re looking for an engineer to join our small, high-impact team responsible for architecting and scaling the core infrastructure behind distributed training pipelines, multimodal data catalogs, and intelligent processing systems that operate over petabytes of data.</p>\n<p>As a software engineer on our data infrastructure team, you&#39;ll design, build, and operate scalable, fault-tolerant infrastructure for LLM Research: distributed compute, data orchestration, and storage across modalities. You&#39;ll develop high-throughput systems for data ingestion, processing, and transformation , including training data catalogs, deduplication, quality checks, and search. You&#39;ll also build systems for traceability, reproducibility, and robust quality control at every stage of the data lifecycle.</p>\n<p>You&#39;ll collaborate with research teams to unlock new features, improve data quality, and accelerate training cycles. You&#39;ll implement and maintain monitoring and alerting to support platform reliability and performance.</p>\n<p>If you&#39;re excited by distributed systems, large-scale data mining, open-source tools like Spark, Kafka, Beam, Ray, and Delta Lake, and enjoy building from the ground up, we&#39;d love to hear from you.</p>\n<p style=\"margin-top:24px;font-size:13px;color:#666;\">XML job scraping automation by <a href=\"https://yubhub.co\">YubHub</a></p>","url":"https://yubhub.co/jobs/job_9be280f4-cbc","directApply":true,"hiringOrganization":{"@type":"Organization","name":"Thinking Machines Lab","sameAs":"https://thinkingmachines.ai/","logo":"https://logos.yubhub.co/thinkingmachines.ai.png"},"x-apply-url":"https://job-boards.greenhouse.io/thinkingmachines/jobs/5013919008","x-work-arrangement":"onsite","x-experience-level":"entry|mid|senior","x-job-type":"full-time","x-salary-range":"$350,000 - $475,000 USD","x-skills-required":["backend language (Python or Rust)","distributed compute frameworks (Apache Spark or Ray)","cloud infrastructure","data lake architectures","batch and streaming pipelines"],"x-skills-preferred":["Kafka","dbt","Terraform","Airflow","web crawler","deduplication","data mining","search","file formats and storage systems"],"datePosted":"2026-04-18T15:54:00.309Z","jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"San Francisco"}},"employmentType":"FULL_TIME","occupationalCategory":"Engineering","industry":"Technology","skills":"backend language (Python or Rust), distributed compute frameworks (Apache Spark or Ray), cloud infrastructure, data lake architectures, batch and streaming pipelines, Kafka, dbt, Terraform, Airflow, web crawler, deduplication, data mining, search, file formats and storage systems","baseSalary":{"@type":"MonetaryAmount","currency":"USD","value":{"@type":"QuantitativeValue","minValue":350000,"maxValue":475000,"unitText":"YEAR"}}},{"@context":"https://schema.org","@type":"JobPosting","identifier":{"@type":"PropertyValue","name":"YubHub","value":"job_e72a55f7-581"},"title":"Senior Software Engineer, Data Acquisition","description":"<p><strong>Senior Software Engineer, Data Acquisition</strong></p>\n<p><strong>Overview:</strong></p>\n<p>The Data Acquisition team within the Foundations organization at OpenAI is responsible for all aspects of data collection to support our model training operations. Our team manages web crawling and GPTBot services and works closely with Data Processing, Architecture, and Scaling teams. We are looking for a skilled Senior Software Engineer to join our Data Acquisition team.</p>\n<p><strong>Responsibilities:</strong></p>\n<ul>\n<li>Own and lead engineering projects in the area of data acquisition including web crawling, data ingestion, and search.</li>\n<li>Collaborate with other sub-teams, such as Data Processing, Architecture, and Scaling, to ensure smooth data flow and system operability.</li>\n<li>Work closely with the legal team to handle any compliance or data privacy-related matters.</li>\n<li>Develop and deploy highly scalable distributed systems capable of handling petabytes of data.</li>\n<li>Architect and implement algorithms for data indexing and search capabilities.</li>\n<li>Build and maintain backend services for data storage, including work with key-value databases and synchronization.</li>\n<li>Deploy solutions in a Kubernetes Infrastructure-as-Code environment and perform routine system checks.</li>\n<li>Conduct and analyze experiments on data to provide insights into system performance.</li>\n</ul>\n<p><strong>Qualifications:</strong></p>\n<ul>\n<li>BS/MS/PhD in Computer Science or a related field.</li>\n<li>6+ years of industry experience in software development.</li>\n<li>Experience with large web crawlers a plus</li>\n<li>Strong expertise in large stateful distributed systems and data processing.</li>\n<li>Proficiency in Kubernetes, and Infrastructure-as-Code concepts.</li>\n<li>Willingness and enthusiasm for trying new approaches and technologies.</li>\n<li>Ability to handle multiple tasks and adapt to changing priorities.</li>\n<li>Strong communication skills, both written and verbal.</li>\n</ul>\n<p><strong>About OpenAI</strong></p>\n<p>OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.</p>\n<p style=\"margin-top:24px;font-size:13px;color:#666;\">XML job scraping automation by <a href=\"https://yubhub.co\">YubHub</a></p>","url":"https://yubhub.co/jobs/job_e72a55f7-581","directApply":true,"hiringOrganization":{"@type":"Organization","name":"OpenAI","sameAs":"https://jobs.ashbyhq.com","logo":"https://logos.yubhub.co/openai.com.png"},"x-apply-url":"https://jobs.ashbyhq.com/openai/70c63d7a-df6f-48f0-b529-f02221e3dc23","x-work-arrangement":"onsite","x-experience-level":"senior","x-job-type":"full-time","x-salary-range":"$293K – $385K","x-skills-required":["large web crawlers","large stateful distributed systems","data processing","Kubernetes","Infrastructure-as-Code","key-value databases","synchronization"],"x-skills-preferred":["trying new approaches and technologies","handling multiple tasks and adapting to changing priorities","strong communication skills"],"datePosted":"2026-03-06T18:42:53.434Z","jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"San Francisco"}},"employmentType":"FULL_TIME","occupationalCategory":"Engineering","industry":"Technology","skills":"large web crawlers, large stateful distributed systems, data processing, Kubernetes, Infrastructure-as-Code, key-value databases, synchronization, trying new approaches and technologies, handling multiple tasks and adapting to changing priorities, strong communication skills","baseSalary":{"@type":"MonetaryAmount","currency":"USD","value":{"@type":"QuantitativeValue","minValue":293000,"maxValue":385000,"unitText":"YEAR"}}},{"@context":"https://schema.org","@type":"JobPosting","identifier":{"@type":"PropertyValue","name":"YubHub","value":"job_02a33cac-468"},"title":"Software Engineer, Data Acquisition","description":"<p><strong>Software Engineer, Data Acquisition</strong></p>\n<p><strong>Overview:</strong></p>\n<p>The Data Acquisition team within the Foundations organization at OpenAI is responsible for all aspects of data collection to support our model training operations. Our team manages web crawling and GPTBot services and works closely with Data Processing, Architecture, and Scaling teams. We are looking for a skilled Software Engineer to join our Data Acquisition team.</p>\n<p><strong>Responsibilities:</strong></p>\n<ul>\n<li>Own and lead engineering projects in the area of data acquisition including web crawling, data ingestion, and search.</li>\n<li>Collaborate with other sub-teams, such as Data Processing, Architecture, and Scaling, to ensure smooth data flow and system operability.</li>\n<li>Work closely with the legal team to handle any compliance or data privacy-related matters.</li>\n<li>Develop and deploy highly scalable distributed systems capable of handling petabytes of data.</li>\n<li>Architect and implement algorithms for data indexing and search capabilities.</li>\n<li>Build and maintain backend services for data storage, including work with key-value databases and synchronization.</li>\n<li>Deploy solutions in a Kubernetes Infrastructure-as-Code environment and perform routine system checks.</li>\n<li>Conduct and analyze experiments on data to provide insights into system performance.</li>\n</ul>\n<p><strong>Qualifications:</strong></p>\n<ul>\n<li>BS/MS/PhD in Computer Science or a related field.</li>\n<li>4+ years of industry experience in software development.</li>\n<li>Experience with large web crawlers a plus</li>\n<li>Strong expertise in large stateful distributed systems and data processing.</li>\n<li>Proficiency in Kubernetes, and Infrastructure-as-Code concepts.</li>\n<li>Willingness and enthusiasm for trying new approaches and technologies.</li>\n<li>Ability to handle multiple tasks and adapt to changing priorities.</li>\n<li>Strong communication skills, both written and verbal.</li>\n</ul>\n<p><strong>About OpenAI</strong></p>\n<p>OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.</p>\n<p><strong>Compensation:</strong></p>\n<ul>\n<li>$293K – $385K • Offers Equity</li>\n</ul>\n<p>The base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience. If the role is non-exempt, overtime pay will be provided consistent with applicable laws. In addition to the salary range listed above, total compensation also includes generous equity, performance-related bonus(es) for eligible employees, and the following benefits.</p>\n<p><strong>Benefits:</strong></p>\n<ul>\n<li>Medical, dental, and vision insurance for you and your family, with employer contributions to Health Savings Accounts</li>\n<li>Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses (parking and transit)</li>\n<li>401(k) retirement plan with employer match</li>\n<li>Paid parental leave (up to 24 weeks for birth parents and 20 weeks for non-birthing parents), plus paid medical and caregiver leave (up to 8 weeks)</li>\n<li>Paid time off: flexible PTO for exempt employees and up to 15 days annually for non-exempt employees</li>\n<li>13+ paid company holidays, and multiple paid coordinated company office closures throughout the year for focus and recharge, plus paid sick or safe time (1 hour per 30 hours worked, or more, as required by applicable state or local law)</li>\n<li>Mental health and wellness support</li>\n<li>Employer-paid basic life and disability coverage</li>\n<li>Annual learning and development stipend to fuel your professional growth</li>\n<li>Daily meals in our offices, and meal delivery credits as eligible</li>\n<li>Relocation support for eligible employees</li>\n<li>Additional taxable fringe benefits, such as charitable donation matching and wellness stipends, may also be provided.</li>\n</ul>\n<p>More details about our benefits are available to candidates during the hiring process.</p>\n<p style=\"margin-top:24px;font-size:13px;color:#666;\">XML job scraping automation by <a href=\"https://yubhub.co\">YubHub</a></p>","url":"https://yubhub.co/jobs/job_02a33cac-468","directApply":true,"hiringOrganization":{"@type":"Organization","name":"OpenAI","sameAs":"https://jobs.ashbyhq.com","logo":"https://logos.yubhub.co/openai.com.png"},"x-apply-url":"https://jobs.ashbyhq.com/openai/41d9d129-2e58-4ad3-be81-2e5096f4da4d","x-work-arrangement":"onsite","x-experience-level":"mid","x-job-type":"full-time","x-salary-range":"$293K – $385K • Offers Equity","x-skills-required":["large web crawlers","large stateful distributed systems","data processing","Kubernetes","Infrastructure-as-Code concepts"],"x-skills-preferred":["trying new approaches and technologies","handling multiple tasks and adapting to changing priorities","strong communication skills"],"datePosted":"2026-03-06T18:39:38.574Z","jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"San Francisco"}},"employmentType":"FULL_TIME","occupationalCategory":"Engineering","industry":"Technology","skills":"large web crawlers, large stateful distributed systems, data processing, Kubernetes, Infrastructure-as-Code concepts, trying new approaches and technologies, handling multiple tasks and adapting to changing priorities, strong communication skills","baseSalary":{"@type":"MonetaryAmount","currency":"USD","value":{"@type":"QuantitativeValue","minValue":293000,"maxValue":385000,"unitText":"YEAR"}}},{"@context":"https://schema.org","@type":"JobPosting","identifier":{"@type":"PropertyValue","name":"YubHub","value":"job_0e1c6ab7-b1a"},"title":"Backend Software Engineer - Search, Crawler Team","description":"<p>We are seeking an experienced Backend Software Engineer to join our Crawler team. In this role, you will design, develop, and operate systems that ingest, process, and manage web-scale data in support of our next generation of advanced search technologies.</p>\n<p><strong>What you&#39;ll do</strong></p>\n<p>In this role, you will take ownership of and lead projects focused on developing large-scale web crawlers, ingestion pipelines, and data processing systems. You will build, maintain, and optimize core backend and frontend components for crawler services, including storage, retrieval, and UI dashboards for data management.</p>\n<ul>\n<li>Take ownership of and lead projects focused on developing large-scale web crawlers, ingestion pipelines, and data processing systems.</li>\n<li>Build, maintain, and optimize core backend and frontend components for crawler services, including storage, retrieval, and UI dashboards for data management.</li>\n</ul>\n<p><strong>What you need</strong></p>\n<ul>\n<li>Minimum of 5 years of software development experience, with strong knowledge of data structures and algorithms in at least one of the following languages: Python, C++, Rust, or Go.</li>\n<li>Experience with large-scale web crawlers is highly desirable.</li>\n<li>Proven experience building, deploying, and optimizing high-load, distributed, and hardware-adjacent services.</li>\n<li>Deep understanding of cloud infrastructure, with hands-on experience in Kubernetes (K8s) and AWS.</li>\n<li>Demonstrated passion for writing clean, efficient, and scalable systems.</li>\n</ul>\n<p style=\"margin-top:24px;font-size:13px;color:#666;\">XML job scraping automation by <a href=\"https://yubhub.co\">YubHub</a></p>","url":"https://yubhub.co/jobs/job_0e1c6ab7-b1a","directApply":true,"hiringOrganization":{"@type":"Organization","name":"Perplexity","sameAs":"https://www.perplexity.ai/","logo":"https://logos.yubhub.co/perplexity.ai.png"},"x-apply-url":"https://jobs.ashbyhq.com/perplexity/94ccf41e-d3e1-41aa-9569-c3bcbffc4184","x-work-arrangement":"onsite","x-experience-level":"senior","x-job-type":"full-time","x-salary-range":null,"x-skills-required":["data structures and algorithms","large-scale web crawlers","cloud infrastructure","Kubernetes (K8s) and AWS"],"x-skills-preferred":["Python","C++","Rust","Go"],"datePosted":"2026-03-04T12:27:29.321Z","jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Belgrade, Berlin, London"}},"employmentType":"FULL_TIME","occupationalCategory":"Engineering","industry":"Technology","skills":"data structures and algorithms, large-scale web crawlers, cloud infrastructure, Kubernetes (K8s) and AWS, Python, C++, Rust, Go"}]}