Showing posts with label Sampling. Show all posts
Showing posts with label Sampling. Show all posts
Thursday, 18 September 2014
Save the dates! Upcoming tweet chats
There have been lots of interesting discussions and topics floating around about new social media in the social scienes. What better way to share than to host some tweet chats! See below for the dates and the topic we will cover for each tweet chat. All times are London time.
Tuesday 7th October, 2014 at 5pm: Representativeness of online samples
Including: What are the geographical inequalities in contributions across different social media platforms? What approaches can we take to address this? How can we weight twitter data? How can we learn about demographics of people on social media, such as age, gender, employment?
Monday 17th November at 5pm: Ready, set, research!: accessing funds and data
Including: You have an idea for a study, how do you go about funding it? What funding streams are available? What are the regulations/restrictions of accessing different streams? How do we get our hands on big data sets from the likes of Google and Twitter?
Tuesday 9th December at 5pm: The changing role of researchers of SM
Including: How is social media changing our identities as researchers, as people? How does this effect our work? How does this impact the field of social sciences?
Remember to include #NSMNSS in all your posts to help us capture all of the discussion. We will provide a transcript of the Tweetchat on our blog following the event.
Friday, 28 February 2014
Using “Small Data” to Improve the Use of “Big Data”

This post was first published on Survey Post on Feb. 3rd, 2014.
Recently, I attended two statistical events in the Washington, DC, area: one was the 23rd Morris Hansen Lecture on “Envisioning the 2030 U.S. Census”; the other was the SAMSI workshop on “Computational Methods for Censuses and Surveys.” “Big data” was a popular keyword at both events and stirred up discussions on how to utilize it (such as from administrative records and online data sources) for current government statistics, especially when combining big data with traditional survey data.
Statisticians are exploring new ways in which big data can be used. The US Census has initiated investigations on using administrative records in the 2020 Census. The National Center for Health Statistics (NCHS) has identified some research opportunities combining multiple data sources. University-based researchers have launched studies on the use of Google trends and other online data in small area estimation.
When big data dominated the mainstream discussion at these events, I started thinking more about “small data.” Can small data help us make better use of big data? Here are some of my thoughts.
- Applying a conventional sampling-based approach to big data: more and more administrative records are collected electronically. Statisticians are excited about using these records that may contain information from the entire population for analytic purposes. Literature in the past two decades has extensively discussed the advantages of administrative records. Processing administrative records data, however, can be quite time consuming. In addition, it can be cumbersome to run analyses on these large datasets because of the large data volume. Especially, when analysts use conventional statistical software, such as SAS, Stata and R, it becomes increasingly complex to handle, store and analyze these data. The question is: is there a way to reduce the data volume and increase computational speed? Applying conventional sampling-based approach (e.g. optimal sampling, calibration weighting) may make a big data smaller and more manageable while allowing researchers to maintain decent data quality.
- Combining non-probability sample data with probability sample data: many big data, such as data collected by Google/Twitter/Facebook, are not census (population) data. We may treat them as non-probability sample data. Elements are chosen arbitrarily in these datasets and there is no way to estimate the probability that each element in the population will be included. Also, it is not guaranteed that each element has a chance of being included, making it impossible either to assess the validity (always measured in terms of “bias”) and reality (always measured in terms of “variance”) of the data. One solution to make the data more representative of the entire population is to combine them with probability sample data (e.g. survey data), which can be relatively smaller. This method can also assist us estimating sample variability and identifying potential bias in big data.
- Using high-quality small data for measuring and adjusting errors in big data: big data is not only non-representative of the target population, but also carry loads of measurement errors because the construct behind a particular measure in these data can differ from the construct that analysts require. To evaluate errors in the big data and improve precision, small survey data can be collected for validation. Take the National Health Interview Survey (NHIS) as an example. This is a household interview survey with only self-reported data. To improve on analyses of the NHIS self-reported data, an imputation-based strategy for using clinical information from an examination-based health survey (i.e. National Health Nutrition Examination Survey, NHANES) was implemented that predicts clinical values from self-reported values and covariates. Estimates of health measures based on the multiply imputed clinical values are different from those based on the NHIS self-reported data alone and have smaller estimated standard errors than those based solely on the NHANES clinical data. Similarly, we may assess potential errors in big data through a more sophisticated and accurate small survey.
While big data provides us massive and timely information from various sources (e.g. social media, administrative records, small data is simple, easy to collect and process, and can be more accurate and representative. Can small data help you when dealing with your big data problems?
Dan Liao is a research statistician at RTI International. She currently works on multiple aspects of data processing and analysis for large, multistage surveys of health care in the United States, including sampling design, calibration weighting, data editing and imputation, statistical disclosure control, and the analysis of survey data. Her survey research interests include multiphase survey designs, combining survey and administrative data, domain estimation, calibration weighting, and regression diagnostics for complex survey data. Dan has a PhD in Survey Methodology from the Joint Program in Survey Methodology at University of Maryland and has published research focusing on regression diagnostics, calibration weighting and predictive modeling.
Tuesday, 29 January 2013
Calling all participants! Recruitment through LinkedIn
Wilma Garvin is currently undertaking research into organisation development and the role of women in business in the 21st century and lectures at the Docklands Business School, University of East London.
My research was to review approaches to organization development in multinationals and government departments in order to create specific case studies. The method used was to undertake face to face and telephone interviews. Sampling using probability sampling seemed a logical approach since the objective was to interview Heads of Organisation Development (OD) or similar i.e. senior people responsible for Organisation Development.
Linkedin was used as the social media. I use Linkedin on a regular basis and I am a member of several OD networks. It seemed like the ideal way to identify and make contact with potential participants. Therefore, the opportunity was already being in OD networks and having contacts either in OD or with contacts in their networks in OD. The challenge was using these in order to find potential participants. These methods of making contact with people through Linkedin were as follows:
- Posting a message with some details of the research in the group area and asking people to contact me
- Asking direct contacts to make an introduce me to a specific person in their network
In using Linkedin, as one source of participants the challenges were:
Ethics: while people voluntarily post their business details on Linkedin, it did seem as if it might be seen as a low level form of stalking by using Linkedin to search for potential participants. Waskul and Douglas (1996 p131) have identified online interaction as neither public or private but as the ‘privately public’ and the ‘publicly private’.
Another challenge was the expectation that there would be a positive response from each of the people in my network to introducing me to their contacts. Some contacts took action immediately even though the contact was on a 3rd level. In a few instances, my contact did not respond and took no action. Where people did try to introduce me to their contacts, the contact declined to participate.
Sampling: Using Linkedin for the sampling frame might be seen as valid although only those who have been a conscious decision to be on Linkedin will be found. This means that the list will be incomplete as it will not represent the whole community of OD specialists.
Linkedin is a public environment and so the people there have chosen to present certain information publicly and so this overcomes some of the ethical issues. It was only being used as a way to contact experts in the field and therefore there would be informed consent. The message posted on the discussion boards provided enough detail and then further details were provided as follow up and in advance of the interviews. However, posting a message on the discussion boards might have been seen as an inappropriate use of the community (Eysenbach and Till 2001).
With regard to ethics, there was also the personal feeling that looking at people’s profile before asking a contact to make an introduction seemed like invasion of privacy may be the change from the ‘old’ attitudes to the new public environment of the internet.
With regard to sampling, it seemed as if self- selection non-probability sampling had to be accepted. On the positive side, it might have been difficult to identify a database of OD specialists since the job titles of OD practitioners can be very different and so using Linkedin does provide a way to identify who these are by drawing on knowledge of the job titles used and the access to Linkedin groups.
The resources used were the knowledge of OD, contacts and specialist groups and the knowledge of using social media.
Linkedin and other social media platforms provide a way to make contact with professionals in specific categories who are willing to take place in research. While self-selection does take place, this can be seen as a positive aspect since the people who did volunteer were passionate about their subject area and were keen to share but also saw it as a way to learn from others.
References
Bakardjieva, M. and Feenberg, A (2001) Involving the virtual subject, Ethics Information Technology, Vol 2:4, p233-240
Eysenbach, G. and Till, J.E. (2001) Ethical issues in qualitative research on internet communities, British Medical Journal, Vol. 323, p1103-5
Waskul, D. and Douglass, M. (1996) Considering the electronic participant: some polemical observations on ethics of on-line research, The Information Society, Vol 12:2, p129-139
Linkedin was used as the social media. I use Linkedin on a regular basis and I am a member of several OD networks. It seemed like the ideal way to identify and make contact with potential participants. Therefore, the opportunity was already being in OD networks and having contacts either in OD or with contacts in their networks in OD. The challenge was using these in order to find potential participants. These methods of making contact with people through Linkedin were as follows:
- Posting a message with some details of the research in the group area and asking people to contact me
- Asking direct contacts to make an introduce me to a specific person in their network
In using Linkedin, as one source of participants the challenges were:
Ethics: while people voluntarily post their business details on Linkedin, it did seem as if it might be seen as a low level form of stalking by using Linkedin to search for potential participants. Waskul and Douglas (1996 p131) have identified online interaction as neither public or private but as the ‘privately public’ and the ‘publicly private’.
Another challenge was the expectation that there would be a positive response from each of the people in my network to introducing me to their contacts. Some contacts took action immediately even though the contact was on a 3rd level. In a few instances, my contact did not respond and took no action. Where people did try to introduce me to their contacts, the contact declined to participate.
Sampling: Using Linkedin for the sampling frame might be seen as valid although only those who have been a conscious decision to be on Linkedin will be found. This means that the list will be incomplete as it will not represent the whole community of OD specialists.
Linkedin is a public environment and so the people there have chosen to present certain information publicly and so this overcomes some of the ethical issues. It was only being used as a way to contact experts in the field and therefore there would be informed consent. The message posted on the discussion boards provided enough detail and then further details were provided as follow up and in advance of the interviews. However, posting a message on the discussion boards might have been seen as an inappropriate use of the community (Eysenbach and Till 2001).
With regard to ethics, there was also the personal feeling that looking at people’s profile before asking a contact to make an introduction seemed like invasion of privacy may be the change from the ‘old’ attitudes to the new public environment of the internet.
With regard to sampling, it seemed as if self- selection non-probability sampling had to be accepted. On the positive side, it might have been difficult to identify a database of OD specialists since the job titles of OD practitioners can be very different and so using Linkedin does provide a way to identify who these are by drawing on knowledge of the job titles used and the access to Linkedin groups.
The resources used were the knowledge of OD, contacts and specialist groups and the knowledge of using social media.
Linkedin and other social media platforms provide a way to make contact with professionals in specific categories who are willing to take place in research. While self-selection does take place, this can be seen as a positive aspect since the people who did volunteer were passionate about their subject area and were keen to share but also saw it as a way to learn from others.
References
Bakardjieva, M. and Feenberg, A (2001) Involving the virtual subject, Ethics Information Technology, Vol 2:4, p233-240
Eysenbach, G. and Till, J.E. (2001) Ethical issues in qualitative research on internet communities, British Medical Journal, Vol. 323, p1103-5
Waskul, D. and Douglass, M. (1996) Considering the electronic participant: some polemical observations on ethics of on-line research, The Information Society, Vol 12:2, p129-139
Friday, 25 January 2013
Videos from our last Knowledge Exchange Seminar
During our last Knowledge Exchange Seminar we looked at three key issues for quantitative researchers using new social media:
Big Data
Populations and sampling
Data visualisation
We asked some of our participants to give a short presentation to open up the discussions on these issues. We filmed each of our presenters and you can find each of their presentations in the links below:
Panos Panagiotopoulos
Carl Miller
Luke Sloan
Patty Kostkova
Scott Hale
Grant Blank
Ralph Schroeder
Big Data
Populations and sampling
Data visualisation
We asked some of our participants to give a short presentation to open up the discussions on these issues. We filmed each of our presenters and you can find each of their presentations in the links below:
Panos Panagiotopoulos
Carl Miller
Luke Sloan
Patty Kostkova
Scott Hale
Grant Blank
Ralph Schroeder
Thursday, 18 October 2012
KES 2 - Populations and sampling
The second session at the Knowledge Exchange Seminar on quantitative
methods on the 26th of September was from Grant Blank of the OII.
The topic was Populations and Sampling and
he asked the questions; What is the “population” on social media platforms? How
do platforms differ in population characteristics? How can we select cases or
sample on social media?
One of the key
issues in terms of sampling online is that it’s difficult to develop a sampling
frame; Grant pointed out that a biased
sampling frame was unavoidable in much online research. However, despite
the potential problems, the advantages of online data collection often outweigh
the challenges, not least because it’s cheap and fast.
Since Twitter
data are so easy to collect, much of the discussion following the session was
around the challenges in sampling Tweets.
How can we get a random representative sample of tweets, especially if we’re
interested in looking at more than just a snapshot of time? It seems to me that
a potential aim for the network might be to put together some guidelines around sampling from Twitter
for new researchers who are looking for guidance. Again, the question was
raised about what kinds of questions Twitter data can really help us to answer,
if we know that Twitter users are not representative of the whole population
and that even getting a random, representative sample of tweets is problematic.
Some case studies and examples of research questions where Twitter data has
been used to good effect could also be helpful to network members.
Little time
was spent discussing sampling from other social media platforms, but an
interesting reference was provided for Gjoka et al (2010) which promotes a
Random Walk technique to obtain an unbiased sample of social network sites,
see:
Subscribe to:
Posts (Atom)