Browse Data & ML Jobs
Search 150 curated tech job listings scraped in real-time from LinkedIn, Glassdoor, RemoteOK, and more. Filter by role, location, seniority, and source to find your next opportunity.
150 jobs found for "Spark"
Page 1 of 8
Lead Data Engineer (Python, AWS, Spark, Kafka, SQL, Snowflake, Databricks, GenAI)
…Preferred Qualifications: Master's Degree 7+ years of experience in application development including Python, SQL, Spark, ETL tools, or AWS Glue 4+ years of experience with a public cloud (AWS, Microsoft Azure … Google Cloud) 4+ years experience with Distributed data/computing tools (MapReduce, Hadoop, Hive, EMR, Kafka, Spark, or MySQL) 4+ year experience working on real-time data and streaming applications 4+ years of experience…
Lead Data Engineer (Python, AWS, Spark, Kafka, SQL, Snowflake, Databricks, GenAI)
…Preferred Qualifications: Master's Degree 7+ years of experience in application development including Python, SQL, Spark, ETL tools, or AWS Glue 4+ years of experience with a public cloud (AWS, Microsoft Azure … Google Cloud) 4+ years experience with Distributed data/computing tools (MapReduce, Hadoop, Hive, EMR, Kafka, Spark, or MySQL) 4+ year experience working on real-time data and streaming applications 4+ years of experience…
Lead Data Engineer (Python, Scala, Spark)
…Azure, Google Cloud) 4+ years experience with Distributed data/computing tools (MapReduce, Hadoop, Hive, EMR, Kafka, Spark, Gurobi, or MySQL) 4+ year experience working on real-time data and streaming applications 4+ years…
Senior Frontend Developer (+ Equity) at AI-native startup backed by YC and Spark Capital
…personal AI assistant for knowledge workers. They've raised $30M in seed funding from Spark Capital, Felicis, and Y Combinator, marking one of the largest seed rounds in YC history. The team…
Python/PySpark Data Engineer
…platforms and intelligent data solutions. The ideal candidate will have strong expertise in Python, Apache Spark, cloud data technologies, data engineering best practices, and the integration of AI/ML capabilities into enterprise data … ETL/ELT workflows for structured, semi-structured, and unstructured data. Develop data processing solutions using Apache Spark, Spark SQL, and DataFrames. Integrate data from APIs, databases, files, streaming platforms, and cloud services. Implement…
Senior Developer - Python / PySpark
…market, trade, risk, and reference data. Develop high-performance data processing solutions using Apache Spark and Azure Databricks. Design and implement reliable data ingestion frameworks for internal and external data sources. Develop … reusable data transformation, validation, and reconciliation frameworks. Optimize Spark applications for scalability, reliability, and performance. Implement data quality controls, monitoring, and governance best practices. Troubleshoot production issues and perform root cause analysis…
Senior Software Engineer - Data Platform
…specialized software engineering role focused on data infrastructure. You will work on systems such as Spark and Databricks infrastructure, Delta Lake on S3, data replication from primary data stores such … operational use cases. Improve the reliability, observability, scalability, security, and developer experience of Samsara’s Spark and Databricks-based data processing platform. Develop internal libraries, APIs, frameworks, and tooling in languages such…
Data Engineer
…based applications, scripts, APIs, and automation solutions. Build and optimize data processing pipelines using PySpark/Apache Spark . Work with relational and NoSQL databases for data ingestion, transformation, and storage. Implement data quality, validation … APIs . Strong understanding of ETL/ELT, data warehousing, data modeling, and database concepts . Experience with Apache Spark and distributed data processing. Experience with at least one cloud platform: AWS, Azure, or GCP . Knowledge…
Sr. Solutions Engineer New
…public cloud platform (AWS, Azure, or GCP) Working knowledge of distributed data systems: Apache Spark™, Delta Lake, or equivalent (Hadoop, Kafka, Flink) Experience leading technical customer conversations — discovery, whiteboarding, architecture reviews Familiarity … Have: Databricks certification or experience with the Databricks Platform Experience with Unity Catalog, Lakeflow Spark Declarative Pipelines, or MLflow Background at a data/AI company, cloud provider, or technical consulting firm Interview Process…
Python Hadoop Engineer
…distributed data processing jobs. Practical experience working with distributed data tools such as Apache Spark, Databricks, or Hadoop ecosystems. Experience building and managing datasets in relational and/or cloud-based data platforms (Teradata … observability for data pipelines (logs, metrics, health checks). Advanced experience with performance tuning of SQL, Spark, or distributed data workflows. Knowledge of data security practices (encryption, masking, PII handling). Experience supporting analytical…
Sr. Solutions Engineer
…public cloud platform (AWS, Azure, or GCP) Working knowledge of distributed data systems: Apache Spark™, Delta Lake, or equivalent (Hadoop, Kafka, Flink) Experience leading technical customer conversations — discovery, whiteboarding, architecture reviews Familiarity … Have: Databricks certification or experience with the Databricks Platform Experience with Unity Catalog, Lakeflow Spark Declarative Pipelines, or MLflow Background at a data/AI company, cloud provider, or technical consulting firm Interview Process…
Senior Data & Software Engineer
…design patterns. What you'll need : Minimum of 5 years' experience with the following: Apache Spark & PySpark Using orchestration tools to deploy data pipelines, including configuring and updating Spark Jobs Advanced Python…
Forward Deployed Engineer
…Azure, GCP) with expertise in at least one Deep experience with distributed computing with Apache Spark™ and knowledge of Spark runtime internals Familiarity with CI/CD for production deployments Working knowledge of MLOps…
Data Scientist II
…PyTorch to third party LLM APIs Process massive amounts of data with Python, SQL and Spark Align with stakeholders through written and verbal communications methods on the approaches and results of projects … Hands-on experience building ML pipelines and working with distributed data processing frameworks like Apache Spark, Databricks, or similar. Intermediate level in at least three of these fields: classification algorithms, natural language…
Data Scientist II
…PyTorch to third party LLM APIs Process massive amounts of data with Python, SQL and Spark Align with stakeholders through written and verbal communications methods on the approaches and results of projects … Hands-on experience building ML pipelines and working with distributed data processing frameworks like Apache Spark, Databricks, or similar. Intermediate level in at least three of these fields: classification algorithms, natural language…
Lead Data Engineer (Python, AWS, SQL, GenAI) (Enterprise Platforms Technology)
…Azure, Google Cloud) Preferred Qualifications: 7+ years of experience in application development including Python, SQL, Spark, ETL tools, AWS Glue 4+ years of experience with a public cloud (AWS, Microsoft Azure, Google … Cloud) 4+ years of experience with Distributed data/computing tools (MapReduce, Hadoop, Hive, EMR, Kafka, Spark, or MySQL) 4+ years of experience working on real-time data and streaming applications 4+ years…
Data Scientist II
…PyTorch to third party LLM APIs Process massive amounts of data with Python, SQL and Spark Align with stakeholders through written and verbal communications methods on the approaches and results of projects … Hands-on experience building ML pipelines and working with distributed data processing frameworks like Apache Spark, Databricks, or similar. Intermediate level in at least three of these fields: classification algorithms, natural language…
Data Scientist II
…PyTorch to third party LLM APIs Process massive amounts of data with Python, SQL and Spark Align with stakeholders through written and verbal communications methods on the approaches and results of projects … Hands-on experience building ML pipelines and working with distributed data processing frameworks like Apache Spark, Databricks, or similar. Intermediate level in at least three of these fields: classification algorithms, natural language…
Data Scientist II
…PyTorch to third party LLM APIs Process massive amounts of data with Python, SQL and Spark Align with stakeholders through written and verbal communications methods on the approaches and results of projects … Hands-on experience building ML pipelines and working with distributed data processing frameworks like Apache Spark, Databricks, or similar. Intermediate level in at least three of these fields: classification algorithms, natural language…
Data Scientist II
…PyTorch to third party LLM APIs Process massive amounts of data with Python, SQL and Spark Align with stakeholders through written and verbal communications methods on the approaches and results of projects … Hands-on experience building ML pipelines and working with distributed data processing frameworks like Apache Spark, Databricks, or similar. Intermediate level in at least three of these fields: classification algorithms, natural language…