back to home

Machine-Learning Engineer Internship at M3

Machine Learning Data Quality Recommendation

This was a two-day machine-learning engineer internship at M3, which runs a large online service for doctors in Japan. We were given raw logs and an open goal on one of their services, so choosing the problem was part of the task. I narrowed it to one group of users who were already active on the site but had never used the feature, and tried to personalize what each of them was shown.

Most of the two days went into checking what my numbers actually measured. The labels I first extracted looked rich, but the counts per item were wildly uneven, and that turned out to be because many rows came from mass re-sends rather than real choices. I discarded 86% of the extracted labels, and the signal became clearer, not weaker. I also found two leaks in my own pipeline, and one artifact I had introduced myself, which I caught because a stress test that should have lowered the score did not. Fixing the evaluation reversed which method looked better.

My conclusion was modest. I could not show that machine learning was needed, because a simple rule did about as well, so I recommended starting with rules and keeping machine learning for the cases the rule cannot handle. Two days are not enough to get accuracy, but they are enough to narrow the problem and to find out what a metric is really measuring, and here that changed my answer.