Header Section
Brief Description
Build, manage and maintain data pipelines.
Qualifications
Training Specialization
Certifications
Job Responsibilities
- Build and schedule ETL pipelines to collect data from source systems, clean and standardize data (process and store error records, handle duplicates, create shared data), assess data quality, aggregate dimension tables, fact tables, OLAP tables, export data to other systems according to detailed design documents.
- Build data transfer pipelines between large data clusters.
- Build processes to clean old data or compress data
- Build data backup processes
- Perform bug fixes identified during development and deployment.
- Identify root causes and fix errors caused by individuals during development and deployment.
- Write documentation, prepare upgrade scripts with upgrade requirements
- Execute upgrade requirements according to existing procedures and scripts.
Interview Questions
- Proficiency in one of the big data storage, processing frameworks or libraries (Hadoop, Spark, Kafka, Nifi)
- Knowledge of database types (RDBMS, Graph Databases, NoSQL Products, ...)
- Solid knowledge of data structures and algorithms:
+ Detailed understanding of basic data types (Integer, Boolean...) and arrays
+ Clear understanding of the relationship between data structures and algorithms
+ Understanding, evaluating complexity and implementing algorithms, for example sorting algorithms: bubble sort, selection sort.., search algorithms - Knowledge of programming, data structures & algorithms
- Proficiency in one programming language (Java, Scala, ...),
- Proficient SQL skills
- Proficiency in one database type (Hive, Oracle, Neo4j, HBase, Cassandra, MongoDB, ..)
- Ability to use log analysis tools to identify root causes of errors.