Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

Thursday, 7 July 2016

Using Big Data to Solve Social Science Problems

Curtis Jessop is a Senior Researcher at NatCen Social Research and is the Network Lead for the NSMNSS network

On Wednesday 29th June I attended a roundtable hosted by our network partners SAGE on using big data to solve social science problems. It was a great day, with contributions from leading researchers and lots of discussion of some of the key issues of working with big data in social science.

Jane Elliott began with an overview of the ESRC’s Big Data Network. She identified the difficulties with data access that earlier phases had faced, but also highlighted key challenges that big data social science currently faces:

1. Methodological
  • Can we apply the same qualitative techniques/statistical inferences we have in the past?
  • Are social scientists (falling) behind in using machine learning & algorithms? What are the implications of these methods?
2. Relevance of research
  • Making sure we use big data to answer pertinent social science questions, and not just focus on methods
3. Ethics at a macro & micro level
  • Working ethically with big data - data security, anonymity, informed consent & data ownership
  • What are the implications of a ‘big data society’/algorithm-led decision making?

New methods, tools and techniques for big data research


Giuseppe Veltri outlined how data-driven science differs from ‘traditional’ social science research as it generates hypotheses and insights from the data, rather than theory, combining abductive, inductive & deductive approaches. Further, Phillip Brooker identified a tension in big data analysis between wanting to use qualitative research approaches with data of a scale that requires numerical treatment. As a result, social scientists need to work with ‘unfamiliar’ techniques and software.


Tools for Big Data analysis


It was generally agreed that existing software are not fit for addressing academic/social science research questions. Also, tools offered by commercial companies are often ‘black boxes’, when social scientists need to be transparent on the algorithms they use as they are part of the methodology.

Many at the roundtable have therefore developed their own tools (e.g. COSMOS, TextonicsChorus, & Method52 from CASM) to enable them to conduct analysis in a manner they wanted to. However, it was felt there was still some way to go - many of these tools are ‘in-house’ and ongoing funding/support is needed to develop something more stable, well-supported, and ‘outward facing’.

Interdisciplinary working


One approach to addressing the challenges of big data analysis is working in interdisciplinary teams (in particular linking between social & computer science departments). Luke Sloan and Mark Carrigan identified the key challenge of this at a ‘human level’ is ensuring a common understanding of language, after which it was easy to have an open discussion and there were rarely disagreements. Mark argued that what was key was not necessarily making sure that everyone had the same definitions, but that there was an understanding that different fields may have different perspectives.

Mark Kennedy, based on his experiences at the Data Science Institute, emphasised the importance of ‘getting excited’ about the right research question, not just focusing on the technology, and then building a team based on what skills you need to fill that gap.

However, attendees felt that there were structural barriers to interdisciplinary working in academia – departmental silos, geography, navigating different funding bodies, finding journals to publish in, and demonstrating value for the REF were all recognised as problems, although it was also mentioned that funding increasingly supported this approach.

Training in the social sciences


Quite early in the discussion, the question was raised that if there is such a clear skills gap in the social sciences, why had universities not responded to it?

Although it was accepted that training needed to address big data methods, there were differing opinions on how feasible this might be. Adding new techniques into methods courses was welcomed, but to what extent was this achievable when these are already packed covering ‘traditional’ methods? Further, given the relative rarity of established social scientists with this skill-set, who would provide this teaching?

Although it was felt that new students are open to using Python or R/new statistical techniques, this scarcity of trainers with the skills to teach both programming and its application within social sciences was again identified as a problem. Giving students (and academics) access to data science training materials that are framed by social science problems, and relevant dummy data to work with, was suggested as a way to start addressing this.

Answering social science questions with Big Data


While discussing his own research, Slava Mikhaylov highlighted that a good way to make impact is, rather than starting with a research question, to aim to solve a problem. This was echoed by Carl Miller, who outlined some principles that Demos follow for making impact:
  • Look beyond academic funders – if research is funded by a government department, they’re going to have to listen to it!
  • Ask the right question – what is interesting to a researcher vs. a policy maker
  • Answer quickly – policy interests change, and research won’t make an impact if everyone’s moved on
  • Diversify outputs – can they be real-time, interactive, engaging?
  • Networking – who are the champions of big data research?

 Carl emphasised that was just the approach that Demos used, and may not be appropriate for all research or audiences. He also mentioned you need to work hard in a new discipline to be responsible and transparent about what your research doesn’t do or say.

Ethics of research using Big Data


Anne Alexander differentiated between the ethics of research using big data and the ethics of doing research in a networked world.

On the latter, Anne felt that there has not been enough reflection on the implications of the ‘datafication’ of human interaction, and that we need to de-mystify these processes and consider what the use of machine learning/algorithms means for society (e.g. their potential for discrimination).

Anne emphasised the need to take into consideration the public’s views on this when considering Big Data research, a point re-enforced by Steve Ginnis, whose work at Ipsos Mori on developing ethical guidelines for social media research drew on public ethics, existing industry guidelines and legal frameworks.

Steve’s research identified that the public both have low awareness of, and are not keen on, their social media data being used for research. This was not just due to concerns about privacy/anonymization – people were uncomfortable with being profiled and its possible implications.

That said, participants were willing to weigh up the risks and benefits, and context (who is doing the research and why) was important. Nonetheless, the ‘fundamentals’ (consent, what information, anonymization, etc.) played a much larger role in whether they felt research using social data was appropriate.

Both Anne & Steve emphasised that ethics is an ongoing process, not a one-off event at the start of a project – they need to be considered at the collection, analysis and publication stages of the research cycle.

Some concluding thoughts


Carl Miller identified that in the context of pressure for evidence-based policy, digital by default, and the open data initiative, there has never been a better time for social scientists to make impact with big data research.

Wednesday’s session demonstrated how far big data analysis in the social sciences has come over recent years and it is impressive to hear how much work has been put into developing the tools and methods to mould this rich, but novel, form of data into social insights.

However, the session also showed that there are number of areas that still need to be addressed if we are to make the most of big data:
  • Access to large data sets continues to be an issue, be they proprietary, public, or administrative. We need to bargain collectively to talk to large, often global, actors and argue for academic access.
  • There is a skills gap among social scientists for analysing big data, and support is needed to help develop the required methodological and programming skills.
  • The interdisciplinary working required for big data analysis can be challenging, and we need to work to enable effective collaboration.
  • Developing an ethical approach to big data analysis is challenging given its novelty, variety, and changing nature. Any framework needs to provide practical guidance to researchers while remaining flexible and responsive to changing contexts.
  • Available tools for big data analysis can be expensive, lack transparency, or inappropriate for social science research. A maintained central library of available tools, with appropriate documentation and guidance could be extremely useful.

Wednesday, 19 November 2014

Making the most out of big data: computer mediated methods

Patrick Readshaw is a Media and Cultural Studies Doctoral Candidate at Canterbury Christ Church University. Patrick is interested in social media as an alternative and empowering source of information on current events, free from the constraints of other agenda-setting media forms. You can contact Patrick by email on p.j.readshaw68@canterbury.ac.uk  

When I was asked to write a blog for NSMNSS, I was certainly excited and being my first post of this kind I was suitably anxious about the prospect. However, my ongoing thesis has never ceased to provide interesting discussions with individuals in linked or parallel fields relating to social media. The main caveat in these discussions is that I often have to try not to over complicate things. With that in mind and my ham-fisted introduction out of the way I want to take some time to break down the value of so called “new media systems” like Twitter and the how I personally go about dealing with the data I collect. 

Since Social Media sites such as “Facebook” burst onto the scene 10 years ago, researchers and market analysts have been looking for a way to tap into the content on these sites. In recent years, there have been several attempts to do this with some being more successful than others (Lewis, Zamith & Hermida, 2013), particularly with regards to the scale of the medium in question. For those uninitiated (apologies to those that are) the term “Big Data” is the catch-all for the enormous trails of information generated by consumers going about their day in an increasingly digitized world (Manyika et al., 2011). It is this sheer volume of information that poses the first hurdle to be overcome when conducting research online. For example, earlier this year I was collecting data on the European Parliamentary Election and generated over 16,000 tweets in about three weeks. Bearing in mind that on average a tweet contains approximately 12 words in 1.5 sentences (Twitter, 2013), for those three weeks I had 196,500 words or 24,500 sentences to come to terms with. That is a lot of data for one person to deal with alone, especially if only applying manual techniques such as content analysis. 

So ultimately you have to ask two questions. Firstly how many undergraduates/interns chained to computers running basic content analysis is it going to take to complete the analysis in a reasonable space of time and whether that analysis is going to be reliable between the analysts. Secondly, while computational methods save time on analysis can you guarantee the same level of depth as with manual content analysis? Considering that content analysis goes beyond basic frequency statistics which can be collected simply from Twitter’s own search engine, I advocate the use of computer mediate techniques in which the data collected can firstly be reduced using filters to removes reTweets or spam responses and secondly to apply hierarchical cluster analysis among others to structure the data somewhat, or at least conceptualise it along a number of important factors. Both Howard (2011) and Papacharissi (2010) utilise this mixed methods approach as do Lewis, Zamith and Hermida (2013) whose method I adapted to my own work and applied as described above. Furthermore these individual pieces of research suggest the value of the medium overall as a source of data, due to its role as one of the primary news disseminators when access to mainstream news media is blocked such as during 2011 Arab Spring events. Burgess and Bruns (2012) have conducted addition research looking at the 2010 federal election campaign in Australia, advising the use of computational methods to reduce their sample to facilitate manual methods ultimately, maintaining depth during content analysis. As can be imagined Lewis, Zamith and Hermida (2013) and Manovich (2012) both support the methodologies utilized by the studies above and advocate making the most of the technical advances that have allowed for the content in question to be organized and harnessed in an efficient way.  

The application of mixed methodologies will continue to develop the techniques integral to facilitating the oncoming age of computational social science (Lazer et al., 2009) or “New Social Science”. While this is the case it is vitally important that while using this readily available source of data is not exploited in a way that could be potentially damaging to the medium as a whole and maintaining good research practice concerning the ethics associated with consumer privacy. As a final aside I would like to remind everyone that this data is hugely fascinating and rich beyond all belief but there are dangers associated with quantifying social life and if possible this should be at front of our minds before, during and after conducting research online (Boyd & Crawford, 2012; Oboler, Welsh & Cruz, 2012).


References

Boyd, d. & Crawford, K. (2012). Critical questions for Big Data: Provocations for a cultural, technological, and scholarly phenomenon. Information, Communication & Society, 15 (5), 662–679.

Burgess, J., & Bruns, A. (2012). (Not) the Twitter election: The dynamics of the #ausvotes conversation in relation to the Australian media ecology. Journalism Practice, 6 (3), 384– 402.
Howard, P. (2011). The digital origins of dictatorship and democracy: Information technology and political Islam. London, UK: Oxford University Press.

Lazer, D., Pentland, A., Adamic, L., Aral, S., Barbási, A., Brewer, D., Christakis, N., Contractor, N., Fowler, J., Gutmann, M., Jebara, T., King, G., Macy, M., Roy, D. & Van Alstyne, M. (2009). Life in the network: The coming age of computational social science. Science, 323 (5915), 721-723.

 Lewis, S. C., Zamith, R., & Hermida, A. (2013). Content Analysis in an Era of Big Data: A Hybrid Approach to Computational and Manual Methods. Journal of Broadcasting & Electronic Media, 57 (1), 34–52.

Manovich, L. (2012). Trending: The promises and the challenges of big social data. In M. K. Gold (Ed.), Debates in the Digital Humanities (pp. 460–475). Minneapolis, MN: University of Minnesota Press.

Manyika, J., Chui, M., Brown, B., Bughin, J., Dobbs, R., Roxburgh, C., & Byers, A. H. (2011). Big data: The next frontier for innovation, competition, and productivity. McKinsey Global Institute.

Oboler, A., Welsh, K., & Cruz, L. (2012). The danger of big data: Social media as computational social science. First Monday, 17 (7-2). Retrieved from
http://firstmonday.org/htbin/cgiwrap/bin/ojs/index.php/fm/article/view/3993/3269.

Papacharissi, Z. (2010). A private sphere: Democracy in a digital age. Cambridge, England: Polity Press.




Thursday, 13 November 2014

The changing nature of who produces and owns data: How will it impact survey research?

Brian Head is a research methodologist at RTI International. This post first appeared on SurveyPost on 20 May, 2014. You can follow Brian on Twitter @BrianFHead.

Cloud Photo

Survey researchers have become interested in big data because it offers potential solutions to problems we’re experiencing with traditional methods. Much of the focus so far has been on social media (e.g., Tweets), but sensors (wearable tech) and the internet of things (IoT) are producing an increasingly rich, complex, and massive source of data. These new data sources could lead to an important change in how individuals see the data collected about them, and thus have ramifications for those interested in gathering and analyzing those data.

Who compiles data?

Quantitative data about people have been gathered for millennia. But with technological advances and identification of new purposes for it, the past 100 years have seen significant increases in the amount of data produced and collected—e.g., data on consumer patterns and other market research, probability surveys, etc.

Common to these data are three factors: 1) the data are a commodity compiled, used, or traded by third parties; 2) generally there are no direct benefits to individuals about whom data are gathered; and 3) the organizations interested in the data gather, store, and analyze it. All this is not to say that throughout history individuals haven’t collected information about themselves. Individuals have collected qualitative data in the form of diaries and biographies. And, they have collected some quantitative data but this has generally to satisfy a third-party (e.g., collecting financial information to file taxes). But, now in addition to all of the data others compile about them, new technologies like wearable technologies (sensors) and IoT devices allow people to voluntarily produce and compile massive amounts of data about themselves and doing so can have a direct benefit to them. (Involuntary data collection through connected devices is already taking place—e.g., internet connected devices are being used for geo-targeting advertising).

Who owns or controls data?

Data are collected in different ways. Census data are collected periodically (intervals vary by nation) through a mandatory government data collection. Surveys generally operate under the requirement of voluntary participation, although there are exceptions.  Much of the consumer data gathered now is done surreptitiously. Examples include browser cookies that collect information about the websites we visit, search engines that collect information about the internet searches people conduct, email providers that scan emails, and apps that use geodata to market goods and services to prospective clients.

It seems the public is increasingly aware of and concerned with the sum of these data collections. According to a recent Robert Wood Johnson Foundation (RWJF) study large majorities of self-tracking app/device users think (84%) they do or want (75%) to own data that are collected with the device. There have been attempts to limit data collection, such as the recent attempt to limit the data the U.S. government collects on citizens.  Advocates of efforts like this tend to cite concerns over burden and privacy. The exponential growth of data collected both voluntarily and involuntarily through apps, sensors, and the IoT may cause similar (perhaps successful) attempts to change government and corporate policies to provide individuals more control over their data. In fact, market researchers are already beginning to respond to such an interest among consumers by offering to pay consumers for access to their browsing history, social network activity, and transactions they conduct online while at the same time giving those consumers control over which data they sell to the brokers.
As the amount of data collected about us increases, there’s a good chance individuals will increasingly see their data as their own, understand the value it has to various third parties, demand more control over it, and to be compensated for it. At first brush that may seem concerning. However, the type of compensation individuals’ desire for data will likely depend on how data will be used. For example, consumers are likely to continue to trade data for convenience in services (see thesis # 12). And, the RWJF report cited above suggests the usual leverages used to gain survey participation—e.g., topic salience and altruism—may work in gaining access to big data when the purpose of the study is for “public good research.”

Need for further research

Further research is needed in this area of big data to answer questions like: 1) to what extent, and how soon, will a larger proportion of the population begin to voluntarily use sensor and IoT devices; 2) will the general public continue to tolerate involuntary data collection when those data are collected by connected devices; 3) will the general public have opinions similar to early adopters in the RWJF about sharing personal data from connected devices with survey researchers; 4) will the leverages that work for gaining survey participation work for gaining access to personal big data or will new/additional leverages be needed; 5) will we be able to use techniques similar to those used to access administrative record data or will we need to develop new protocol for seeking permission to access these data? I look forward to seeing and contributing toward the research to answer these questions. What are your thoughts?

Thursday, 6 November 2014

You Are What You Tweet: An Exploration of Tweets as an Auxiliary Data Source

Ashley Richards is a survey methodologist at RTI International. This post first appeared on SurveyPost on 29, July 2014. 

Last fall at MAPOR , Joe Murphy presented the findings of a fun study he did with our colleague, Justin Landwehr, and me. We asked survey respondents if we could look at their recent Tweets and combine them with their survey data. We took a subset of those respondents and masked their responses on six categorical variables. We then had three human coders and a machine algorithm try to predict the masked responses by reviewing the respondents’ Tweets and guessing how they would have responded on the survey. The coders looked for any clues in the Tweets, while the algorithm used a subset of Tweets and survey responses to find patterns in the way words were used. We found that both the humans and machine were better than random in predicting values of most of the variables.

We recently took this research a step further and compared the accuracy of these approaches to multiple imputation, with the help of our colleague Darryl Creel. Imputation is the approach traditionally used to account for missing data and we wanted to see how the nontraditional approaches stack up. Furthermore, we wanted to check out these approaches because imputation cannot be used in the case where survey questions are not asked. This commonly occurs because of space limitations, the desire to reduce respondent burden, or other factors. I will be presenting on this research at the upcoming Joint Statistical Meetings (JSM), in early August. I’ll give a brief summary here, but if you’d like more details on it please check out my presentation or email me for a copy of the paper.

Income was the only variable for which imputation was the most accurate approach, but the differences between imputation and the other approaches were not statistically significant. Imputation correctly predicted income 32% of the time, compared to 25% for human coders and 26% for the machine algorithm. Considering that there were four income categories and a person would have a 25% chance of randomly selecting the correct response, I am unimpressed with these success rates of 25%-32%.

Human coders outperformed imputation on the other demographic items (age and sex), but imputation was more accurate than the machine algorithm. For these variables, the human coders picked up on clues in respondents’ Tweets. I was one of the coders and found myself jumping to conclusions, but I did so with a pretty good rate of success. For instance, if a Tweeter said “haha” a lot or used smiley faces, I was more likely to guess the person was young and/or female. These are tendencies that I’ve observed personally but I’ve read about them too.

As a coder I struggled to predict respondents’ health and depression statuses, and this was evident in the results. Imputation was better than humans at predicting these, but the machine algorithm was even more accurate. The machine was also best at predicting who respondents voted for in the previous presidential election, with human coders in second place and imputation in last place. As a coder I found that predicting voting was fairly simple among the subset of respondents who Tweeted about politics. Many Tweeters avoided the subject altogether, but those who Tweeted about politics tended to make it obvious who they supported.

twitter_predictions
So what does this all mean? We found that even with a small set of respondents, Tweets can be used to produce estimates with accuracy in the same range or better[1] as imputation procedures. There is quite a bit of room for improvement in our methods that could make them even more accurate. For example, we could use a larger sample of Tweets to train the machine algorithm and we could select human coders who are especially perceptive and detail-oriented. The finding that Tweets are as good or better as imputation is important because imputation cannot be used in the case where survey questions were not asked.

As interesting as these findings may be, they need to be taken with a grain of salt, especially because of our small sample size (n=29).[2] Relying on Twitter data is challenging because many respondents are not on Twitter, and those who are on Twitter are not representative of the general population and may not be willing to share their Tweets for these purposes. Another challenge is the variation in Tweet content. For example, as I mentioned earlier, some people Tweet their political views while others stay away from the topic on Twitter.

Despite these limitations, Twitter may represent an important resource for estimating values that are desired but not asked for in a survey. Many of our survey respondents are dropping clues about these values across the Internet, and now it’s time to decide if and how to use them. How many clues have you dropped about yourself online? Is your online identity revealing of your true characteristics?!?

[1] Even if approaches using Tweets may be more accurate than imputation, they require more time and money and in many cases may not be worth the tradeoff. As discussed later, these findings need to be taken with a grain of salt.

[2] We had more than 2,000 respondents, but our sample size for this portion of the study was greatly reduced after excluding respondents who don’t use Twitter, respondents who did not authorize our use of their Tweets, and respondents whose Tweets were not in English. Furthermore, half of the remaining respondents’ Tweets were used to train the machine algorithm.

Thursday, 18 September 2014

Save the dates! Upcoming tweet chats


There have been lots of interesting discussions and topics floating around about new social media in the social scienes. What better way to share than to host some tweet chats! See below for the dates and the topic we will cover for each tweet chat. All times are London time.


Tuesday 7th October, 2014 at 5pm: Representativeness of online samples

Including: What are the geographical inequalities in contributions across different social media platforms? What approaches can we take to address this? How can we weight twitter data? How can we learn about demographics of people on social media, such as age, gender, employment?


Monday 17th November at 5pm: Ready, set, research!: accessing funds and data

Including: You have an idea for a study, how do you go about funding it? What funding streams are available? What are the regulations/restrictions of accessing different streams? How do we get our hands on big data sets from the likes of Google and Twitter?


Tuesday 9th December at 5pm: The changing role of researchers of SM

Including: How is social media changing our identities as researchers, as people? How does this effect our work? How does this impact the field of social sciences?


Remember to include #NSMNSS in all your posts to help us capture all of the discussion. We will provide a transcript of the Tweetchat on our blog following the event.

Thursday, 8 May 2014

Understanding Geert Lovink’s book “Networks without a cause: A critique of social media”: A video by Akin Olaniyan

Akin Olaniyan is a student in the MA in Social Media at the University of Westminster

Watch the video interview here! https://www.youtube.com/watch?v=GV-PBy3iaDk

He sounds very much like a techno-pessimist. When I first read Professor Geert Lovink’s book, ‘Networks without a cause: A critique of social media,’ the first thing that struck me was the sense of despair that runs through all the chapters. With chapters devoted to Facebook and the crisis of identity, big data and the ‘Googlisation of our lives,’ and a proposition to divorce the study of social media from media studies, there appeared to be no other way to understand Geert Lovink. And that was the starting point of my interview with him conducted via Skype. As you can see in the video (https://www.youtube.com/watch?v=GV-PBy3iaDk), when I asked him what was the source of the despair I felt running through his book, Geert describes his frustration that the Internet has become too centralized and cites cloud computing as evidence. He feels it was time the Internet went back to the original format of small networks.

One of the comments on the blurb of his book, by McKenzie Wark (author of Gamer Theory and Professor of Culture and Media at The New School), describes Geert as our Tin Tin. ‘Like canny adventurer, he travels the world discovering new frontiers of both folly and invention,’ McKenzie writes and that was instantly confirmed when Geert – after he found where I was from – expressed an interest in working in Nigeria if there was a good chance to teach critical Internet Cultures.

I find that this interview is a rare insight into the minds and works of a scholar who sounds techno-pessimistic but whose research and works focus on making the Internet more workable.

On big data and social media alternatives for instance, Geert, whose works have focused on developing alternative social media, is critical of big data and the concentration of power in the hands of few corporations like Google.  “When we talk of alternatives to social media we refer to alternative variations of the known platforms. So, an alternative to Google would be a search engine without all this commercial bias, a search engine that will be based on other algorithmic principles,” he says. And on Facebook, Geert says: “When we talk of alternatives to Facebook, we mean a social network that is truly local and doesn’t work with this ridiculous notion of friends as a general principle of connecting people.” He sounded to me more like an activist than a university teacher.

When asked how to best arrange funding for such alternative social media platforms, Geert admits there may be no global solution owing to differences from country to country but mentioned subscription and the public library model as possible options for consideration.

I finally asked him whether he considers himself as belonging to the same school of thought as Evgeny Morozov and Neil Postman since I found as I felt his arguments sound techno-pessimistic. He appeared to smile at the question but looked serious as he answered: “I strongly believe that we have to team up with start-ups, programmers and insert a political and cultural agenda there. I come from a background where we see ourselves as developers of the network and that is different from traditional techno pessimist point of view. I have absolute joy in the constructive value of ruthless radical critique.” Overall, I find Geert an interesting scholar, especially after I received copies of other books he co-authored, which he freely offered. A quick look at two of those books, ‘Unlike Us Reader: Social Media Monopolies and their Alternatives,’ and ‘Video Vortex Reader: Responses to YouTube,’ confirm my feeling that this is no ordinary scholar but one who is at the same time an activist.


Wednesday, 5 March 2014

Keeping up with technology: What is “scientific lag” and can we proactively reduce it?

In 2011 then Census Director Robert Groves wrote on the Census Director’s Blog about the burgeoning volume of “organic data”—data that, as opposed to “designed data,” have no meaning until they are used (surveys are a primary example of the latter). He noted that finding ways to combine these two types of data to increase the “information-to-data ratio” was a challenge, but also represented the future of surveys. Using terms identified as “big data descriptors” in Groves’ piece, as well as a few other terms I think qualify, I put together the graph below to show the number of AAPOR presentation titles between 2010 and 2013 that contain a big data descriptor.1,2
big data descriptors in AAPOR presentation titles 2010-2013
One take away is the increased interest researchers have shown in big data over the past few years. An equally important lesson is that almost all of the attention big data has received from AAPOR members—at least measured by the number of presentations they’ve done—has been on social networking sites (SNS). I found only one presentation in the past four AAPOR conference programs that contained a big data descriptor for a non-social media topic—a demonstration in 2012 by Ben Waber on the use of wearable sensors for measuring behavior.
To some extent this is explained by scientific lag. Just like there is cultural lag—the time between the emergence of a new technology and when culture catches up—there is a lag time between when consumers adopt technologies and when our research methodologies catch up (i.e., scientific lag = cultural lag + time until research methodologies using those technologies are implemented). And, technologies often don’t remain static, but rather evolve making it a continuous game of catchup (development) for research methodologists. I’ll go into more detail about this in a presentation I’m giving at the AAPOR conference this year, but one quick example from the annals of survey research history is the development of computer-assisted telephone interviewing (CATI). While telephone exchanges had existed for almost a century and programmable computers emerged in the 1940s, it took until 1971 for CATI systems to be developed by market researchers, another five years for academic researchers to begin using it, and the federal government another seven years to implement its use. Certainly, cultural lag played a role. It took years for enough households to have telephones for probability based telephone sampling to make sense. In addition, it took time for the programmable computer to develop into a device usable for this purpose. But, it also took researchers time to figure out such a system was possible and the value it presented.
Now, let’s fast forward a bit. In 1997, one of the first SNSs, sixdegrees.com, was created.  It lasted until 2001. A host of other networking sites, the ones most of us are familiar with, sprang up in the early 2000s—Myspace (2003), Facebook (2004), and Twitter (2006). There are, I suppose, two ways of looking at the cultural and scientific lags and SNSs. On the one hand, it took a few years SNSs to grow to significant numbers. For example, it took Facebook four years (2004 – 2008) to grow to 100 million users. Within four years of that development there were multiple presentations at AAPOR on the subject. That’s certainly much faster than the development of CATI technology/adoption. On the other hand, social researchers took nearly a decade from the birth of widely popular SNSs to begin formally recognizing their research utility.
Now, we may be at the cusp of another such tsunami of consumer technology adoption. Groups disagree on the exact timing (e.g., Forbes says 2014 and MIT Technology Review says 2013), but the evidence points to the start of rapid growth in the use of internet connected sensors and devices for a multitude of purposes. I’ve recently written about how and why I think the devices and the IoT will affect social science data collection.
My question is whether the research community can be more proactive, and therefore decrease the scientific lag between adoption and research implementation. My hope is we will and that it will have a positive effect on survey data collection.
I’ll be presenting more thoughts on this topic at AAPOR and look forward to the discussion we have about big data in the session. Between now and then I’d welcome the thoughts others have on or experiences others have had with using wearable tech, sensors, or the IoT for research.
This was first posted on Survey Post  on 24/02/14
Brian Head is a research methodologist at RTI International with 5 years of experience in the government and not-for-profit research sectors.  Training in sociology and research methods and statistics led him to a career in research where his work has included questionnaire design and evaluation,  managing data collection efforts, and qualitative and quantitative data analysis.
Share


    Friday, 28 February 2014

    Using “Small Data” to Improve the Use of “Big Data”

    Digital Globe
    This post was first published on Survey Post on Feb. 3rd, 2014.
    Recently, I attended two statistical events in the Washington, DC, area: one was the 23rd Morris Hansen Lecture  on “Envisioning the 2030 U.S. Census”; the other was the SAMSI workshop on “Computational Methods for Censuses and Surveys.” “Big data” was a popular keyword at both events and stirred up discussions on how to utilize it (such as from administrative records and online data sources) for current government statistics, especially when combining big data with  traditional survey data.
    Statisticians are exploring new ways in which big data can be used. The US Census has initiated investigations on using administrative records in the 2020 Census. The National Center for Health Statistics (NCHS) has identified some research opportunities combining multiple data sources. University-based researchers  have launched studies on the use of Google trends and other online data in small area estimation.
    When big data dominated the mainstream discussion at these events, I started thinking more about “small data.” Can small data help us make better use of big data? Here are some of my thoughts.
    1. Applying a conventional sampling-based approach to big data: more and more administrative records are collected electronically. Statisticians are excited about using these records that may contain information from the entire population for analytic purposes. Literature in the past two decades has extensively discussed the advantages of administrative records. Processing administrative records data, however, can be quite time consuming. In addition, it can be cumbersome to run analyses on these large datasets because of the large data volume. Especially, when analysts use conventional statistical software, such as SAS, Stata and R, it becomes increasingly complex to handle, store and analyze these data. The question is: is there a way to reduce the data volume and increase computational speed? Applying conventional sampling-based approach (e.g. optimal sampling, calibration weighting) may make a big data smaller and more manageable while allowing researchers to maintain decent data quality.
    2. Combining non-probability sample data with probability sample data: many big data, such as data collected by Google/Twitter/Facebook, are not census (population) data. We may treat them as non-probability sample data.  Elements are chosen arbitrarily in these datasets and there is no way to estimate the probability that each element in the population will be included. Also, it is not guaranteed that each element has a chance of being included, making it impossible either to assess the validity (always measured in terms of “bias”) and reality (always measured in terms of “variance”) of the data. One solution to make the data more representative of the entire population is to combine them with probability sample data (e.g. survey data), which can be relatively smaller. This method can also assist us estimating sample variability and identifying potential bias in big data.
    3. Using high-quality small data for measuring and adjusting errors in big data: big data is not only non-representative of the target population, but also carry loads of measurement errors because the construct behind a particular measure in these data can differ from the construct that analysts require. To evaluate errors in the big data and improve precision, small survey data can be collected for validation. Take the National Health Interview Survey (NHIS) as an example. This is a household interview survey with only self-reported data. To improve on analyses of the NHIS self-reported data, an imputation-based strategy for using clinical information from an examination-based health survey (i.e. National Health Nutrition Examination Survey, NHANES) was implemented that predicts clinical values from self-reported values and covariates. Estimates of health measures based on the multiply imputed clinical values are different from those based on the NHIS self-reported data alone and have smaller estimated standard errors than those based solely on the NHANES clinical data. Similarly, we may assess potential errors in big data through a more sophisticated and accurate small survey.
    While big data provides us massive and timely information from various sources (e.g. social media, administrative records, small data is simple, easy to collect and process, and can be more accurate and representative.  Can small data help you when dealing with your big data problems?


    Dan Liao is a research statistician at RTI International. She currently works on multiple aspects of data processing and  analysis for large, multistage surveys of health care in the United States, including sampling design, calibration weighting, data editing and imputation, statistical disclosure control, and the analysis of survey data. Her survey research interests include multiphase survey designs, combining survey and administrative data, domain estimation, calibration weighting, and regression diagnostics for complex survey data. Dan has a PhD in Survey Methodology from the Joint Program in Survey Methodology at University of Maryland and has published research focusing on regression diagnostics, calibration weighting and predictive modeling.