Job Description
DISA Technologies develops HPSA (High Pressure Slurry Attrition) mineral processing technology. Over several years of lab, pilot, and field testing we've accumulated a large body of test data, including feed characterization, particle size, chemistry, and process performance. This data is spread across many projects, sites, and file formats. Before we can use it in analytics and machine learning, it needs to be found, catalogued, and put into a consistent format.
Together with the ML Engineer, you'll own the consolidation of our historical test data end-to-end, working from an inventory schema, format specification template, and tabular data template we provide. The goal is a governed, catalogued set of clean data files ready to load into our cloud data platform.
What you'll do
- Build a complete inventory of historical data sources: location, client, project, site/program, equipment unit, date range, format, and completeness — extending and correcting an existing inventory rather than starting from scratch
- Write a short format specification for the common record types you encounter, with example files for each
- Transcribe test files into a consistent tabular template and save them to a designated staging area
- Produce a data quality report covering gaps, incomplete or suspect analyses, and contradictions between files
- After the core work: exploratory data analysis to identify candidate correlations across datasets (e.g., feed characteristics vs. process performance) and short feasibility memos on which questions the historical data can and can't support
- Possible additional work depending on progress: applying file naming/tagging conventions to legacy folders, metadata preparation for a document-system migration, cross-referencing test records against related measurement datasets
Requirements
- Currently enrolled in or recently graduated from a program in computer science, data science, statistics, engineering, or a related field
- Comfortable with Python and pandas (or equivalent) for reading, cleaning, and reshaping tabular data
- Experience working with messy real-world files — Excel workbooks, CSVs, PDFs, inconsistent naming and structure
- Careful and methodical; willing to document what you find, including what's missing or doesn't add up
- Able to work independently from a defined schema and templates, and to ask for clarification when a file doesn't fit them
- Clear written communication
Nice to have
- Exposure to SQL, data catalogs, or data lake / "bronze-silver-gold" concepts
- Basic statistics or EDA experience (matplotlib/seaborn, Jupyter)
- Interest in mineral processing, chemistry, or industrial process data. No prior domain knowledge required
- Familiarity with SharePoint/OneDrive at a power-user level
Benefits
We provide a managed laptop or virtual desktop; the role involves handling confidential client data and requires signing a confidentiality acknowledgment