Blog

  • Where’s the API that can tell me that this photo contains a puppy and a can of Coke?

    Where’s the API that can tell me that this photo contains a puppy and a can of Coke?

    puppy and a can of coke

    Photo by Ahmad van der Breggen on Flickr.

    We’ve gotten very good at extracting and disambiguation entities from text data. You can license a commodity system, and there are API and even open source tools that work fairly well.

    However, a large percentage of content that people share is not primarily text (a back-of-the-envelope guess says around 18%), and we currently have very little automated insight into that content.

    I know this is a very hard problem, but I’m continuously surprised by how few people seem to be working on it. Any ideas?


  • Help, I’m the first data scientist at my company!

    Help, I’m the first data scientist at my company!

    I moderated a panel at DataGotham with Adam Laiacano from TumblrFred Benenson from Kickstarter, and Roberto Medri from Etsy about being the first data scientist at a company. We covered everything from what people’s job responsibilities are, the tools they use, successes, failures, how they are integrated into an organization, and how they have hired other data scientists to join them. The panelists were concise, articulate, and intelligent. Watch it below!


  • Hey Yahoo, You’re Optimizing the Wrong Thing

    Hey Yahoo, You’re Optimizing the Wrong Thing

    I was visiting my grandparents yesterday, and my grandfather asked for help e-mailing an article to some of his friends. I asked him to show me how he normally writes an e-mail, and taught him the magic of copy and paste (it is amazing if you haven’t seen it before) but I noticed that in the course of sending an e-mail and checking on his inbox, he clicked on this ad three times.

    When I asked about it, he didn’t realize he had clicked the ad — he just thought these screens popped up randomly — because he didn’t realize that his hands were shaking on the trackpad.

    I’m sure the data says that that’s the optimal place on the screen for the ad. I’m sure tons of people ‘click’ on it. I’m also sure it’s wrong, and it results in a terrible experience.

    It’s common sense, but experiences like this are great reminders that data only takes us so far, and creativity and clear thinking are always required to find the best solutions.

    Yahoo, please fix this!


  • How do you prioritize research?

    How do you prioritize research?

    One of the most fun and challenging parts of my job is setting bitly’s research agenda. We’re a startup, so this means prioritizing the set of questions we look into in the context of what will be most beneficial for the rest of the business, for the short and long-term, by creating opportunity and opening up potential futures. We work on a wide variety of projects, from pure research to press collaborations to infrastructure and experimental products.

    We always have a list of research questions way longer than we have time and resources to pursue, so we developed a process for evaluating whether a given question is worth pursuing at a particular time.

    This is the kind of process that I’ve only discussed with several people over whisky (thanks!), but not seen written up. I initially had a much longer list of questions but have decided to keep it as simple as possible, to frame a discussion but not dictate or burden it. I hope it’s helpful and I would love to hear about other appproaches.

    For each research question that we might look into, we ask the following:

    1. State the research question.
    2. How do we know when we’ve won?
    3. Assume we’ve solved this question perfectly. What are the first things that we’ll build with it?
    4. If everyone in the world uses this, how does it change human behavior?
    5. What’s the most evil thing that can be done with this?

    State the research question.

    It’s important to state the question in language that everyone can understand. The bitly team comes from a variety of scientific and business backgrounds, and we’ve developed some of our own common vocabulary, but it still takes a bit of effort to make sure that everyone understand the fundamental challenge and why it’s interesting.

    How do we know when we’ve won?

    Here we define the metrics that we’ll use to measure our success. For some questions, this is obvious, and for others it’s impossible to define — we can at least acknowledge that ahead of time.

    Assume we’ve solved this question perfectly. What are the first things that we’ll build with it?

    This question allow us to assess the potential business and product impact. What capabilities will we have with this that we don’t have now? It allows us to keep the long-term research vision in mind while still optimizing for shorter-term opportunities.

    If everyone in the world uses this, how does it change human behavior?

    What’s the maximum potential impact of this work? If it’s not inspiring, is it worth pursuing at all?

    What’s the most evil thing that can be done with this?

    I don’t ask this question to encourage evil (>:]) but as a creative tool for expanding how we think about validity, impact, and potential applications of the research. The label evil is so ridiculous that it permits people share their craziest ideas. Plus, it’s always a fun conversation to have.

    Finally

    I’m always revising this list, and I would love to hear how you think about prioritizing your work.


  • DataGotham: The Empire State of Data

    DataGotham: The Empire State of Data

    I’m extremely excited about DataGotham, a conference that I’m co-hosting with friends and fellow New York data nerds Drew, John, and Mike.

    DataGotham is a celebration of the NYC data community, and will bring together professionals from all industries in New York that are built around data, from finance to fashion and from startups to the Fortune 500 and government. The event is September 13th – 14th at NYU, with tutorials and The Great Data Extravaganza Show (with cocktails!) at the Tribeca Rooftop Thursday evening, and a single track conference Friday. Our speakers and sponsors are all amazing. You can register now.

    While DataGotham is definitely a labor of love, there are numerous reasons to do it. I believe that New York has a distinct data philosophy — the study of human behavior — that is unique and should be celebrated. We have an large population of local badass data hackers, and our community will only grow stronger if we can build relationships across the industry divides. Finally, there’s an opportunity for all of us to influence the future of data science, and this event will highlight some voices that might not otherwise be heard.

    I hope to see you there!

    (Also, anyone who made it this far through can register with code “dataGothamist” for 25% off :) )