Multi-photo evidence fusion and VLM benchmarking for automated form filling - BC-1038
Project type: ResearchDesired discipline(s): Engineering - computer / electrical, Engineering, Computer science, Mathematical Sciences, Statistics / Actuarial sciences
Company: Anonymous
Project Length: 4 to 6 months
Preferred start date: 11/02/2026
Language requirement: English
Location(s): Richmond / Vancouver, BC, Canada
No. of positions: 1
Desired education level: Master'sPhD
Open to applicants registered at an institution outside of Canada: No
About the company:
The partner organization (ANONYMOUS) is an established Canadian electrical utility consulting firm based in British Columbia, with a team of approximately 95 professionals. The company provides engineering services to major Canadian utilities across transmission, distribution, substation, and telecommunications infrastructure. A significant part of its practice is distribution pole replacement design: assessing end-of-life wood poles and producing the engineering design required to replace them safely and to standard. The company is investing in an applied AI/R&D program that uses computer vision and machine learning to automate the interpretation of field photographs of utility poles, so that engineering design information currently extracted manually can be pre-filled automatically with calibrated confidence scores and clear flags for designer review. Interns joining this program work with a real production dataset of field imagery and design records, alongside the company's internal ML, platform, and engineering teams, with a direct path from research results to deployment in a production design workflow.
Describe the project.:
What is the project about / main goal: Each pole is documented by several photographs — top views, full views, and close-ups. This systems-oriented project fuses predictions across a pole's photo set into a single draft design form with per-field confidence and structured JSON output, and benchmarks whether an end-to-end vision-language model can fill the form directly or whether the existing ensemble of specialist models remains the better production path.
Main tasks to be performed by the candidate:
• Reconcile top-view, full-view, and close-up photos into one pole-level prediction
• Benchmark prompted / fine-tuned VLM form filling against specialist model outputs
• Produce structured JSON output mapped to the design-form schema, with confidence and review flags
• Compare accuracy, cost, latency, calibration, and operational reliability
Methodology/techniques to be used: Multi-view aggregation and VLM evaluation; target: beat best single-photo accuracy on ≥ 80% of fields, with a clear production cost/accuracy recommendation.
Required expertise/skills:
VLMs and prompt / fine-tune experimentation; ML systems thinking; evaluation design.
Optional: assets
Production dataset of utility-pole field imagery with linked engineering design records; internal labeling platform and labeling pipeline; existing trained baseline models to build on; cloud GPU compute; day-to-day co-supervision by the company's internal ML and platform team.

