Where’s the API that can tell me that this photo contains a puppy and a can of Coke?

Where’s the API that can tell me that this photo contains a puppy and a can of Coke?

puppy and a can of coke

Photo by Ahmad van der Breggen on Flickr.

We’ve gotten very good at extracting and disambiguation entities from text data. You can license a commodity system, and there are API and even open source tools that work fairly well.

However, a large percentage of content that people share is not primarily text (a back-of-the-envelope guess says around 18%), and we currently have very little automated insight into that content.

I know this is a very hard problem, but I’m continuously surprised by how few people seem to be working on it. Any ideas?