Job Description
Junior Data Engineer
Location: Blouberg, Cape Town
Employment Type: Permanent
Salary: R35,000 – R42,000 per month
Level: Junior
About the Role
An exciting opportunity is available for a Junior Data Engineer to join a growing, data-driven environment based in Blouberg, Cape Town.
The successful candidate will form part of the data team and assist in building, maintaining and improving the data infrastructure that supports business intelligence, reporting, analytics and machine learning initiatives.
This is an excellent opportunity for an ambitious Data Engineer who already has a solid foundation in SQL, data pipelines, ETL/ELT processes and data modelling and is looking to further develop their technical capabilities within a collaborative environment.
You will work closely with Data Scientists, Business Analysts and other technical stakeholders to ensure that reliable, accurate and well-structured data is available across the organisation.
Key Responsibilities
Data Pipeline Development
- Build, maintain and monitor reliable ETL and ELT data pipelines
- Assist with ingesting data from multiple source systems, including databases, APIs and third-party platforms.
- Monitor pipeline performance and identify potential failures or bottlenecks.
- Assist with improving the reliability, scalability and efficiency of existing pipelines.
- Support scheduled and automated data processing workflows.
Data Architecture & Modelling
- Assist with the design and maintenance of data models and database schemas.
- Create and maintain data warehouse tables to support analytics and reporting requirements.
- Work with dimensional modelling concepts, including star and snowflake schemas.
- Assist with structuring data to ensure it is accessible, consistent and suitable for downstream users.
- Support ongoing improvements to the organisation's data architecture.
Data Quality & Testing
- Implement data validation and quality checks across pipelines and datasets.
- Assist with automated testing of data processes.
- Monitor data accuracy, completeness and freshness.
- Investigate discrepancies and identify the source of data-quality issues.
- Help ensure that datasets delivered to downstream users are reliable and fit for purpose.
SQL & Performance Optimisation
- Develop and maintain SQL queries used for data extraction, transformation and reporting.
- Troubleshoot slow-running or inefficient queries.
- Assist with query optimisation and database performance improvements.
- Refactor existing queries and code where required to improve efficiency and maintainability.
Maintenance & Troubleshooting
- Monitor existing pipelines and investigate failures.
- Troubleshoot data integration and processing issues.
- Assist with resolving production data incidents.
- Refactor and improve existing code as the data environment evolves.
- Support ongoing maintenance of the organisation's data infrastructure.
Collaboration
- Work closely with Data Scientists and Business Analysts to understand their data requirements.
- Translate reporting, analytics and machine-learning requirements into reliable datasets.
- Assist stakeholders with understanding available data sources and structures.
- Participate in technical discussions relating to data architecture, pipelines and data quality.
Documentation
- Maintain clear technical documentation for data pipelines and processes.
- Maintain data dictionaries and definitions.
- Document data models, schemas and transformations.
- Record relevant architectural decisions and changes.
- Ensure documentation remains accurate as systems and processes evolve.
Minimum Requirements
- Relevant tertiary qualification in Computer Science, Information Technology, Data Science, Software Engineering, Information Systems or a related discipline would be advantageous.
- Previous practical experience in a Data Engineering, Data Analytics, BI, Database or similar technical environment
- Good working knowledge of SQL
- Understanding of ETL/ELT concepts and data pipelines
- Understanding of relational databases and database structures.
- Exposure to data warehousing concepts.
- Understanding of data modelling principles.
- Familiarity with dimensional modelling, star schemas and/or snowflake schemas
- Ability to work with structured data from multiple sources.
- Basic programming or scripting capability relevant to data engineering.
- Strong analytical and problem-solving ability.
- Comfortable troubleshooting technical and data-related issues.
Highly Advantageous
Exposure to one or more of the following would be beneficial:
- Python
- Advanced SQL
- REST APIs / API integrations
- Cloud-based data platforms
- Data warehouses
- Data orchestration tools
- Git / version control
- Automated data testing
- CI/CD principles
- Business Intelligence and reporting environments
- Machine-learning data requirements
- Database and query performance optimisation
Technical Competencies
The ideal candidate should demonstrate a developing technical foundation across:
Data Engineering: ETL/ELT pipelines, data ingestion, transformation, pipeline monitoring, data integration
Databases: SQL, relational databases, schemas, query optimisation, database fundamentals
Data Warehousing: Data warehouse concepts, dimensional modelling, star schemas, snowflake schemas
Data Quality: Validation, testing, reconciliation, monitoring, accuracy and data freshness
Development: Scripting/programming, debugging, code maintenance, refactoring and version control
Integration: APIs, databases and third-party data sources
Personal Attributes
We are looking for someone who is:
- Analytical and naturally curious
- Detail-oriented and quality focused
- Technically minded
- A logical problem solver
- Comfortable investigating and resolving problems
- Able to work independently while still collaborating effectively with a team
- Willing to learn and develop new technologies
- Organised and disciplined with documentation
- Able to communicate technical information clearly
- Proactive in identifying potential improvements
What Would Make You Stand Out?
You are likely to be particularly well suited to this opportunity if you have already had hands-on exposure to building or maintaining ETL/ELT pipelines, can confidently work with SQL, understand how data warehouses and dimensional models are structured, and enjoy figuring out why data or pipelines are not behaving as expected.
This is a Junior Data Engineer opportunity, so applicants are not expected to know everything. However, candidates should already have a solid technical foundation and be able to demonstrate genuine practical exposure to data engineering concepts.
Location
This position is based in Blouberg, Cape Town Candidates should be based within reasonable travelling distance or be able to reliably commute to the workplace.
Remuneration
R35,000 – R42,000 per month, depending on experience, technical capability and overall suitability for the position.