Multi-photo evidence fusion and VLM benchmarking for automated form filling - BC-1038
Genre de projet: RechercheDiscipline(s) souhaitée(s): Génie - informatique / électrique, Génie, Informatique, Sciences mathématiques, Statistiques / études actuarielles
Entreprise: Anonymous
Durée du projet: 4 à 6 mois
Date souhaitée de début: 11/02/2026
Langue exigée: Anglais
Emplacement(s): Richmond / Vancouver, BC, Canada
Nombre de postes: 1
Niveau de scolarité désiré: MaîtriseDoctorat
Ouvert aux candidatures de personnes inscrites à un établissement à l’extérieur du Canada: No
Au sujet de l’entreprise:
The partner organization (ANONYMOUS) is an established Canadian electrical utility consulting firm based in British Columbia, with a team of approximately 95 professionals. The company provides engineering services to major Canadian utilities across transmission, distribution, substation, and telecommunications infrastructure. A significant part of its practice is distribution pole replacement design: assessing end-of-life wood poles and producing the engineering design required to replace them safely and to standard. The company is investing in an applied AI/R&D program that uses computer vision and machine learning to automate the interpretation of field photographs of utility poles, so that engineering design information currently extracted manually can be pre-filled automatically with calibrated confidence scores and clear flags for designer review. Interns joining this program work with a real production dataset of field imagery and design records, alongside the company's internal ML, platform, and engineering teams, with a direct path from research results to deployment in a production design workflow.
Veuillez décrire le projet.:
What is the project about / main goal: Each pole is documented by several photographs — top views, full views, and close-ups. This systems-oriented project fuses predictions across a pole's photo set into a single draft design form with per-field confidence and structured JSON output, and benchmarks whether an end-to-end vision-language model can fill the form directly or whether the existing ensemble of specialist models remains the better production path.
Main tasks to be performed by the candidate:
• Reconcile top-view, full-view, and close-up photos into one pole-level prediction
• Benchmark prompted / fine-tuned VLM form filling against specialist model outputs
• Produce structured JSON output mapped to the design-form schema, with confidence and review flags
• Compare accuracy, cost, latency, calibration, and operational reliability
Methodology/techniques to be used: Multi-view aggregation and VLM evaluation; target: beat best single-photo accuracy on ≥ 80% of fields, with a clear production cost/accuracy recommendation.
Expertise ou compétences exigées:
VLMs and prompt / fine-tune experimentation; ML systems thinking; evaluation design.
Optional: assets
Production dataset of utility-pole field imagery with linked engineering design records; internal labeling platform and labeling pipeline; existing trained baseline models to build on; cloud GPU compute; day-to-day co-supervision by the company's internal ML and platform team.

