NewGen Technologies Inc.

Data & Software Engineer

NewGen Technologies Inc.  •  Arlington, VA / Chantilly, VA / Herndon, VA / McLean, VA (Onsite)  •  3 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

The Data & Software Engineer works with a small team to build complex data flows for a custom application. Successful candidate will have advanced Python programming skills, familiarity with Java, an understanding of data security, privacy, governance and compliance principles and a demonstrated history of building production data pipelines and ETL workflows at scale.

The candidate must have experience: Building end-to-end data pipelines leveraging Python. Using orchestration tools to deploy data pipelines, including configuring and updating Spark Jobs. Containerizing and deploying applications in cloud environments like AWS. Working with MySQL and PostgreSQL including performance tuning, schema design, and query optimization for complex, analytical workloads. Leveraging industry standard tools for code control (Git, IaaC control, etc.). Working with data catalogs, tracking data lineage and handling a variety of data formats, including Geospatial. Using Bash scripting for automation and data processing tasks. Integrating Al/ML services and models.

Responsibilities
  • Work with stakeholders to understand data requirements, assess feasibility, and design appropriate solutions with minimal oversight
  • Leverage strong problem-solving and debugging skills for data quality issues, pipeline failures, and performance bottlenecks
  • Leverage a background in large-scale data migration or platform modernization efforts
  • Contribute to data engineering documentation, best practices, and design patterns
Requirements
  • TS/SCI FSP Clearance
  • Minimum of 5 years' experience:
    • Demonstrated experience building production data pipelines and ETL/EL workflows at scale
    • Proficiency with Apache Spark and PySpark for distributed data processing
    • Advanced Python programming skills including data manipulation libraries (Pandas, NumPy) and data engineering best practices
    • Understanding of data security, privacy, governance, and compliance principles
    • Experience with workflow orchestration tools (such as Step Functions, Airflow)
    • Familiarity with containerization (such as Docker or Podman) and deploying data applications in cloud environments
    • Experience with AWS services (S3, Lambda, Step Functions)
    • Experience with PostgreSQL and MySQL in production environments, including performance tuning and schema design
    • Demonstrated experience with SQL and query optimization for complex analytical workloads
    • Experience with version control (Git) and Cl/CD practices for data pipelines
    • Demonstrated ability to work with stakeholders to understand data requirements, assess feasibility, and design appropriate solutions with minimal oversight
    • Strong problem-solving and debugging skills for data quality issues, pipeline failures, and performance bottlenecks
Highly Desired Skills
  • Experience with data Lakehouse architecture using Apache Iceberg
  • Hands-on experience configuring, deploying, and integrating data platform components:
    • Apache Ranger (access control and data governance)
    • Trino (distributed SQL query engine)
    • Data catalogs (Unity Catalog OSS, Apache Polaris, etc.)
    • Apache Superset {data visualization and dashboarding)
  • Proficiency with Bash scripting for automation and data processing tasks
  • Experience with Infrastructure as Code (Terraform or CloudFormation) for data infrastructure
  • Familiarity or experience with tracking data lineage and associated tooling such as Open lineage
  • Familiarity or experience with Java
  • Familiarity with data quality frameworks, testing methodologies, and validation strategies
  • Background with large-scale data migrations or platform modernization efforts
  • Experience integrating Al/ML services and models (translation, OCR, speech-to-text, NLP, language detection, topic modeling), LLMs, and RAG {retrieval-augmented generation) pipelines
  • Familiarity with geospatial data processing {H3, PostGIS, or similar)
  • Contributions to data engineering documentation, best practices, and design patterns
  • Experience with NoSQL databases (DynamoDB, etc.)

About Us

For more than 20 years, NewGen Technologies has solved our clients’ toughest IT challenges with integrity, security, and outstanding service by delivering both technology and talent. We have helped secure borders, have used artificial intelligence (AI) to fight terror, aided the identification of criminals, and have helped to prevent crime through the introduction of biometrics. Our team of Highly Cleared Specialists have hard-to-find skills and expertise in a wide spectrum of technologies to provide solutions that transform business processes and solve problems of national significance. #CJ
NewGen Technologies Inc.

About NewGen Technologies Inc.

Welcome to NewGen Technologies, Inc.

NewGen Technologies specializes in developing and implementing solutions to your IT challenges-especially those involving specialized technologies. Founded in 1997, NewGen grew quickly and became immediately successful because of our primary commitment to satisfying our customers.

In a world where IT solutions rely on specialized talents, NewGen Technologies and its team of IT specialists have hard-to-find skills and expertise in a spectrum of specialized technologies. NewGen's mission is to provide you with solutions to your IT challenges with integrity, security, and outstanding service.

Why NewGen?

Our formula for success is simple-we deliver high quality products and services. We develop and maintain a staff of highly trained, experienced consultants who devise and execute creative, effective solutions to our clients' IT challenges.

Industry
Unknown
Company Size
51-200 employees
Headquarters
Fulton, Maryland
Year Founded
1997
Social Media