Artificial Intelligence Engineer
qonnectiq- Posted 2 hours ago
- Be among the first 10 applicants
Job Description
Qonnectiq is a software company. We build systems for organisations whose field and technical teams have produced large archives of daily operational reports, technical summaries, instrument logs, and end-of-project reports. Those archives hold most of what an organisation has learned, but they sit in thousands of PDFs and vendor files nobody can query. We are building a small team to change that, and we value engineers who care about getting the numbers right.
Role Description
This is a full-time, on-site Artificial Intelligence Engineer role based in Jakarta Barat. You will be the first AI hire on this team.
Your mission is to turn a large document archive into a structured, auditable knowledge base, then use it to generate draft technical planning documents that engineers can actually work from. Every figure the system produces must trace back to the report it came from.
The same foundation later supports predictive work, but that is not the first year. Most of the value in year one comes from getting the data right: roughly 60% data engineering, 30% AI and search, 10% testing and tooling.
Job Description
1. Design and own the database schema that acts as the source of truth for the entire system.
2. Implement versioning and as-of-date logic, so the system can answer what a value was based on what we knew at a given date.
3. Make provenance first-class: every value carries its source document, location within it, extraction method, and confidence.
4. Build parsing pipelines for PDFs with complex table layouts, specialised technical data formats, and scanned documents requiring OCR.
5. Extract structured entities: operating parameters, equipment specifications, instrument readings, activity sequences, downtime, and costs.
6. Build entity resolution so documents from different sources and vendors attach to the correct asset, since the source data has no unique identifier.
7. Normalise units and measurement reference systems across sources.
8. Build a validation layer that flags contradictions between sources instead of silently picking one.
9. Deduplicate overlapping work periods across reports, and publish data quality reports non-technical users can read.
10. Model relationships between assets: shared work area, geographic proximity, similar technical conditions, shared equipment and contractors.
11. Build search that filters on structured fields first and then searches by meaning, since most of this data is numeric and tabular.
12. Design how the system assembles a draft document from retrieved data, with every figure traceable to its source.
13. Assess whether we use a commercial AI provider or run models on our own infrastructure, and keep that layer replaceable either way.
14. Build test sets from previously approved documents, and repeatable checks that catch quality regressions before release.
15. Ship with safeguards, monitoring, and human review, and close the feedback loop so corrections from engineers re-enter the knowledge base and improve later results.
Qualifications
- Strong experience designing relational database schemas, including versioning of historical data and handling values that get revised over time. This is the most important qualification for this role.
- At least 5 years of experience as an AI/ML Engineer.
- Hands-on experience building data pipelines: ingestion, transformation, validation, scheduling.
- Experience with document parsing or information extraction from PDFs or other semi-structured formats with inconsistent layouts.
- Experience building and shipping applications on large language models, including retrieval-augmented generation (RAG).
- Experience testing systems that do not return the same answer every time, and building your own way of measuring whether output is good enough.
- Able to reason concretely about the trade-offs between using a commercial AI provider and running models in-house.
- Strong SQL, and familiarity with at least one database that supports similarity search.
- Able to explain technical decisions to non-engineers and work alongside domain specialists.
- Solid technical English. Source documents are in English and dense with industry abbreviations.
Nice To Have
- Predictive modelling on operational data with limited sample sizes.
- Experience running AI models on your own infrastructure rather than through a provider.
- Knowledge graph or graph database experience.
- OCR on low-quality scanned technical documents.
- Systems built under strict data security constraints: on-premise, air-gapped, or data residency requirements.
- Exposure to heavy engineering, energy, manufacturing, or similar industries.
- Working knowledge of Java, enough to read our backend code and debug integration issues.
Please note: this is not a research role, and predictive modelling is not the first year's work. The deliverable is a system engineers actually use, and most of the value comes from getting the data right.
More Info
Key Skills
similarity search
data pipelines
testing systems
large language models
retrieval-augmented generation
document parsing
