View all jobs
Principal Data Engineer – (Hadoop/Big Data, AWS, Python, Kinesis) - Irvine, CA - Onsite - Contract to Hire - AS
Position: Senior Data Engineer – (Hadoop/Big Data, AWS, Python, Kinesis)
Location: Irvine, CA - Onsite role
Position Type: Contract to Hire
Responsibilities
You will be responsible for all aspects of data acquisition, data transformation, analytics scheduling and operationalization to drive high-visibility, cross-division outcomes. Investigate, evaluate, test and recommend technical solutions for future systems.
They will support software developers, database architects, data scientists on data initiatives and will contribute optimal data delivery architecture.
What you will be doing:
Data Operations
Lead the creation of data environments and/or data sets to serve a wide range of data users, including but not limited to Data Scientists, Data Analysts, Business Analysts etc.
Perform offline analysis of large data sets using components of a big data software ecosystem.
Validate the solution of root cause analysis escalated by various technical staff in multiple organizations and with differing levels of expertise.
Investigate, evaluate, test and recommend technical solutions for future systems.
Data Management
Own product data sets from the definition phase through to production deployment (end-to-end).
Provide solutions for the design and implementation of Hadoop EMR Cluster/ Big Data Infrastructure.
Deploy Hadoop/Big Data/Spark and database storage Infrastructures in AWS cloud.
Monitor HDFS/Hadoop/Spark and related software releases, third-party utilities with emphasis on overall system performance.
Lead and develop tools and procedures to monitor and automate system tasks on servers and clusters
Lead and collaborate with other teams to design, develop, and deploy data tools that support both operations and product use cases.
Data Design
Lead and design distributed, scalable, and reliable data pipelines that ingest and process data at scale and in real-time.
Lead and design big data technologies and prototype solutions to improve data processing architecture.
Qualifications
What we are looking for:
Bachelor’s degree in computer science, computer engineering, or a related technical field.
12+ years of professional experience as a data software engineer; or 16+ years of related experience as a data software engineer in lieu of 4-year degree.
2+ years of experience with AWS cloud or other cloud Big Data computing design, provisioning, and tuning.
Related AWS certification, preferred.
Previous experience as a Data Engineer / Database Administrator and/or Business Intelligence Analyst.
Expertise in database concepts, object and data modeling techniques and design principles.
Expertise in database architectures, software, and facilities
Expertise with programming languages - Python (required), Scala, Ruby, R Database technologies - SQL, performance tuning concepts, AWS RDS, RedShift, MySQL
Expertise with big data batch processing tools: Hadoop MapReduce, ElasticSearch, PIG, Hive, Cascading/Scalding, Apache Spark, AWS EMR
Expertise with stream-processing systems: Kinesis, Kafka, MQTT
Expertise with relational NoSQL databases including DyanamoDB
Expertise in writing JSON, XML, YAML and other data definition schemas
Excellent verbal and written communication skills necessary to effectively collaborate in a team environment and present and explain technical information and provide advice to management.
Ability to work on advanced complex technical projects or business issues requiring state of the art technical knowledge or industry.
Ability to work on significant and unique issues where analysis of situations or data requires an evaluation of intangibles.
Exercises independent judgment in methods, techniques and evaluation criteria for obtaining results.
Ability to lead and mentor junior engineers and colleagues.
Location: Irvine, CA - Onsite role
Position Type: Contract to Hire
Responsibilities
You will be responsible for all aspects of data acquisition, data transformation, analytics scheduling and operationalization to drive high-visibility, cross-division outcomes. Investigate, evaluate, test and recommend technical solutions for future systems.
They will support software developers, database architects, data scientists on data initiatives and will contribute optimal data delivery architecture.
What you will be doing:
Data Operations
Lead the creation of data environments and/or data sets to serve a wide range of data users, including but not limited to Data Scientists, Data Analysts, Business Analysts etc.
Perform offline analysis of large data sets using components of a big data software ecosystem.
Validate the solution of root cause analysis escalated by various technical staff in multiple organizations and with differing levels of expertise.
Investigate, evaluate, test and recommend technical solutions for future systems.
Data Management
Own product data sets from the definition phase through to production deployment (end-to-end).
Provide solutions for the design and implementation of Hadoop EMR Cluster/ Big Data Infrastructure.
Deploy Hadoop/Big Data/Spark and database storage Infrastructures in AWS cloud.
Monitor HDFS/Hadoop/Spark and related software releases, third-party utilities with emphasis on overall system performance.
Lead and develop tools and procedures to monitor and automate system tasks on servers and clusters
Lead and collaborate with other teams to design, develop, and deploy data tools that support both operations and product use cases.
Data Design
Lead and design distributed, scalable, and reliable data pipelines that ingest and process data at scale and in real-time.
Lead and design big data technologies and prototype solutions to improve data processing architecture.
Qualifications
What we are looking for:
Bachelor’s degree in computer science, computer engineering, or a related technical field.
12+ years of professional experience as a data software engineer; or 16+ years of related experience as a data software engineer in lieu of 4-year degree.
2+ years of experience with AWS cloud or other cloud Big Data computing design, provisioning, and tuning.
Related AWS certification, preferred.
Previous experience as a Data Engineer / Database Administrator and/or Business Intelligence Analyst.
Expertise in database concepts, object and data modeling techniques and design principles.
Expertise in database architectures, software, and facilities
Expertise with programming languages - Python (required), Scala, Ruby, R Database technologies - SQL, performance tuning concepts, AWS RDS, RedShift, MySQL
Expertise with big data batch processing tools: Hadoop MapReduce, ElasticSearch, PIG, Hive, Cascading/Scalding, Apache Spark, AWS EMR
Expertise with stream-processing systems: Kinesis, Kafka, MQTT
Expertise with relational NoSQL databases including DyanamoDB
Expertise in writing JSON, XML, YAML and other data definition schemas
Excellent verbal and written communication skills necessary to effectively collaborate in a team environment and present and explain technical information and provide advice to management.
Ability to work on advanced complex technical projects or business issues requiring state of the art technical knowledge or industry.
Ability to work on significant and unique issues where analysis of situations or data requires an evaluation of intangibles.
Exercises independent judgment in methods, techniques and evaluation criteria for obtaining results.
Ability to lead and mentor junior engineers and colleagues.