Showing posts with label software. Show all posts
Showing posts with label software. Show all posts

Thursday, 7 July 2016

Using Big Data to Solve Social Science Problems

Curtis Jessop is a Senior Researcher at NatCen Social Research and is the Network Lead for the NSMNSS network

On Wednesday 29th June I attended a roundtable hosted by our network partners SAGE on using big data to solve social science problems. It was a great day, with contributions from leading researchers and lots of discussion of some of the key issues of working with big data in social science.

Jane Elliott began with an overview of the ESRC’s Big Data Network. She identified the difficulties with data access that earlier phases had faced, but also highlighted key challenges that big data social science currently faces:

1. Methodological
  • Can we apply the same qualitative techniques/statistical inferences we have in the past?
  • Are social scientists (falling) behind in using machine learning & algorithms? What are the implications of these methods?
2. Relevance of research
  • Making sure we use big data to answer pertinent social science questions, and not just focus on methods
3. Ethics at a macro & micro level
  • Working ethically with big data - data security, anonymity, informed consent & data ownership
  • What are the implications of a ‘big data society’/algorithm-led decision making?

New methods, tools and techniques for big data research


Giuseppe Veltri outlined how data-driven science differs from ‘traditional’ social science research as it generates hypotheses and insights from the data, rather than theory, combining abductive, inductive & deductive approaches. Further, Phillip Brooker identified a tension in big data analysis between wanting to use qualitative research approaches with data of a scale that requires numerical treatment. As a result, social scientists need to work with ‘unfamiliar’ techniques and software.


Tools for Big Data analysis


It was generally agreed that existing software are not fit for addressing academic/social science research questions. Also, tools offered by commercial companies are often ‘black boxes’, when social scientists need to be transparent on the algorithms they use as they are part of the methodology.

Many at the roundtable have therefore developed their own tools (e.g. COSMOS, TextonicsChorus, & Method52 from CASM) to enable them to conduct analysis in a manner they wanted to. However, it was felt there was still some way to go - many of these tools are ‘in-house’ and ongoing funding/support is needed to develop something more stable, well-supported, and ‘outward facing’.

Interdisciplinary working


One approach to addressing the challenges of big data analysis is working in interdisciplinary teams (in particular linking between social & computer science departments). Luke Sloan and Mark Carrigan identified the key challenge of this at a ‘human level’ is ensuring a common understanding of language, after which it was easy to have an open discussion and there were rarely disagreements. Mark argued that what was key was not necessarily making sure that everyone had the same definitions, but that there was an understanding that different fields may have different perspectives.

Mark Kennedy, based on his experiences at the Data Science Institute, emphasised the importance of ‘getting excited’ about the right research question, not just focusing on the technology, and then building a team based on what skills you need to fill that gap.

However, attendees felt that there were structural barriers to interdisciplinary working in academia – departmental silos, geography, navigating different funding bodies, finding journals to publish in, and demonstrating value for the REF were all recognised as problems, although it was also mentioned that funding increasingly supported this approach.

Training in the social sciences


Quite early in the discussion, the question was raised that if there is such a clear skills gap in the social sciences, why had universities not responded to it?

Although it was accepted that training needed to address big data methods, there were differing opinions on how feasible this might be. Adding new techniques into methods courses was welcomed, but to what extent was this achievable when these are already packed covering ‘traditional’ methods? Further, given the relative rarity of established social scientists with this skill-set, who would provide this teaching?

Although it was felt that new students are open to using Python or R/new statistical techniques, this scarcity of trainers with the skills to teach both programming and its application within social sciences was again identified as a problem. Giving students (and academics) access to data science training materials that are framed by social science problems, and relevant dummy data to work with, was suggested as a way to start addressing this.

Answering social science questions with Big Data


While discussing his own research, Slava Mikhaylov highlighted that a good way to make impact is, rather than starting with a research question, to aim to solve a problem. This was echoed by Carl Miller, who outlined some principles that Demos follow for making impact:
  • Look beyond academic funders – if research is funded by a government department, they’re going to have to listen to it!
  • Ask the right question – what is interesting to a researcher vs. a policy maker
  • Answer quickly – policy interests change, and research won’t make an impact if everyone’s moved on
  • Diversify outputs – can they be real-time, interactive, engaging?
  • Networking – who are the champions of big data research?

 Carl emphasised that was just the approach that Demos used, and may not be appropriate for all research or audiences. He also mentioned you need to work hard in a new discipline to be responsible and transparent about what your research doesn’t do or say.

Ethics of research using Big Data


Anne Alexander differentiated between the ethics of research using big data and the ethics of doing research in a networked world.

On the latter, Anne felt that there has not been enough reflection on the implications of the ‘datafication’ of human interaction, and that we need to de-mystify these processes and consider what the use of machine learning/algorithms means for society (e.g. their potential for discrimination).

Anne emphasised the need to take into consideration the public’s views on this when considering Big Data research, a point re-enforced by Steve Ginnis, whose work at Ipsos Mori on developing ethical guidelines for social media research drew on public ethics, existing industry guidelines and legal frameworks.

Steve’s research identified that the public both have low awareness of, and are not keen on, their social media data being used for research. This was not just due to concerns about privacy/anonymization – people were uncomfortable with being profiled and its possible implications.

That said, participants were willing to weigh up the risks and benefits, and context (who is doing the research and why) was important. Nonetheless, the ‘fundamentals’ (consent, what information, anonymization, etc.) played a much larger role in whether they felt research using social data was appropriate.

Both Anne & Steve emphasised that ethics is an ongoing process, not a one-off event at the start of a project – they need to be considered at the collection, analysis and publication stages of the research cycle.

Some concluding thoughts


Carl Miller identified that in the context of pressure for evidence-based policy, digital by default, and the open data initiative, there has never been a better time for social scientists to make impact with big data research.

Wednesday’s session demonstrated how far big data analysis in the social sciences has come over recent years and it is impressive to hear how much work has been put into developing the tools and methods to mould this rich, but novel, form of data into social insights.

However, the session also showed that there are number of areas that still need to be addressed if we are to make the most of big data:
  • Access to large data sets continues to be an issue, be they proprietary, public, or administrative. We need to bargain collectively to talk to large, often global, actors and argue for academic access.
  • There is a skills gap among social scientists for analysing big data, and support is needed to help develop the required methodological and programming skills.
  • The interdisciplinary working required for big data analysis can be challenging, and we need to work to enable effective collaboration.
  • Developing an ethical approach to big data analysis is challenging given its novelty, variety, and changing nature. Any framework needs to provide practical guidance to researchers while remaining flexible and responsive to changing contexts.
  • Available tools for big data analysis can be expensive, lack transparency, or inappropriate for social science research. A maintained central library of available tools, with appropriate documentation and guidance could be extremely useful.

Monday, 11 August 2014

7 Ways NVivo Helps Researchers Handle Social Media Data

Kathleen McNiff is a blogger with QSR International, the people that brought you NVivo. Get in touch with Kath on Twitter @KMcNiff.


Imagine you’re sitting on a qualitative goldmine—in-depth interviews, focus groups, intriguing survey results, nuanced observations and a comprehensive lit review.
 All the traditional boxes are ticked and yet there’s the nagging feeling that something is missing. 

Chances are, it’s social media—and that’s probably why you’re here.
 
The Challenges
There are so many impassioned and revealing conversations taking place online that it’s becoming harder (and more dangerous) to ignore them. 
But embracing social media is not straightforward and you may be grappling with questions such as:
 
  • How do I build social media into my research design?
  • What platforms are worth concentrating on?
  • How should I collect the data?
  • What tools and methods should I use to analyse it?
NVivo gives you a practical way to face these challenges.
 
NCapture the web
If you already work with NVivo, you’ll know that it’s a tool for organizing and analyzing qualitative data—but you may not realise that NVivo 10 for Windows comes with a raft of features to support your foray into the brave new world of social media.
 
It all starts with NCapture.
 
This small but powerful plugin sits quietly at the top of your browser (Internet Explorer or Chrome) and lets you capture web pages and social media—and then bring them into NVivo for analysis. It’s a bit like that helpful elephant from Evernote.
You can also capture YouTube videos and conversations from Facebook, Twitter or LinkedIn. This is a boon for researchers who want to facilitate ‘online focus groups’ using these social media platforms—as well as for those who want to get a well-rounded view of their topic by following the latest conversations.
 This brief video (with lovely music) shows you how to gather Twitter data using NCapture:
 


If you use NVivo 10 for Mac—stay tuned, because NCapture is coming soon.
Now, let’s focus on 7 ways NVivo helps you to make sense of your social media data.
 
#1: Gather tweets or posts in a dataset
 
You can search for tweets in your browser and then use NCapture to pull them into a PDF or dataset. The dataset it especially handy because you can filter or sort the content—and use tools to slice and dice the data in different ways.


 
You can do the same for discussions and posts from Facebook or LinkedIn—on these platforms, you can also use the biographical data from user profiles to compare attitudes (men vs women, young vs old—that kind of thing).
 
Sometimes you have to work around the limitations of a particular platform. For example, the number of tweets you can capture is determined by Twitter and can vary depending on the vagaries of Twitter traffic. To follow a particular topic over time, the best approach is to take captures at periodic intervals.
If you want to know more about the inner workings of Twitter—there is a fantastic post right here on NSMNSS blog.
#2: Visualize the most frequently used words
 
You can run a Word Frequency query to see which words contributors are using most often—this can help you get a handle on the themes in your social media data.
 
Visualizing the results in a word cloud may spark insights and reveal connections—they can also liven up a presentation, final paper or blog post.


 
#3: Map the location of tweets or posts
 
You can open a map to see where the action is—and then use this as a launching point for further investigation. For example, you could click on a pin to see the tweets or posts from a particular location.

 
 
#4: Chart users by the number of followers
Shares, likes and follows are the new social currency and they can help to inform your research. If you’re exploring Twitter users - you can create a chart to compare the numbers:
#5: Organize the content into themes
 
You know that qualitative goldmine I mentioned earlier? Well, you can bring the whole thing into NVivo 10 for Windows (including your newly NCaptured social media data) and use ‘coding’ to organize it into themes.
 
For example, whenever you see a reference to ‘education’—whether it be in an interview, article or social media conversation—you can select the content and code it at a ‘node’. Then you can open the node (which is a fancy word for container) and explore all the references to ‘education’ in one place.
 
Coding is a great way to wrangle the chaos of qualitative data—and it’s slightly addictive.
#6: Explore by username or hashtag
 
Do you want to gather tweets from a particular user or hashtag? If your tweets are in a dataset, then ‘auto coding’ is your answer. You can easily roll up the tweets to coding collections by username, by hashtag (what did everyone say about #NSMNSS?), or even by location.
 
Speaking of hashtags - why not start your own twitter chat to gather feedback about an issue or idea?
 
#7: Press the ‘Analyze This’ button
 
Can’t find it?
That’s because, as awesome as NVivo is, it won’t do the analysis for you.
 
Don’t fret because you’ll find plenty of tools for querying the data as well as ways of organizing your own analytical insights (including memos, annotations, models and framework matrices).
Explore the possibilities
Social media has opened a Pandora’s box of opportunities for qualitative research—but you needn’t be overwhelmed because NVivo provides a safe to place to put the box while you explore its contents.
 
Maybe you’re already using NVivo to analyze your social media data?
 
 
Share how you are using NCapture in a short blog post by emailing NSMNSS@natcen.ac.uk