Showing posts with label big data. Show all posts
Showing posts with label big data. Show all posts

Thursday, December 21, 2023

One Defence Data Overly Ambitious

The ABC reports that "$100m Defence contract with KPMG rife with governance failures, review finds" (Linton Besser, Andrew Greene, ABC, 20 December 2023). The contact concerned Defence ICT2284 "One Defence Data" (1DD). 1DD is an ambitious project to unify all Defence data. This project appears to have been overly ambitious, and should have been scaled back.

As it happens I was the Senior Policy Adviser on Data Administration Standards at DoD from 1990 to 1994, and know how hard bringing disparate data sources jealously guarded by different stakeholders is. 

1DD includes an enterprise-wide data catalogue.  

"The One Defence Data Program (Program) will establish and deliver the governance, standards and change management to drive information management transformation across the Department of Defence (Defence). Tranche 1 will develop the foundational and technical enablement capabilities that will continue to be built out over future tranches of the Program." ICT 2284 – One Defence Data Program – Tranche 1, KPMG Australia Technology Solutions Pty Limited (KTech), for DoD, Mon 02 May 2022. 

The Department has a Defence Data Strategy 2021-2023:



Thursday, April 28, 2022

Software for Fast Storage Hardware With Prof Willy Zwaenepoel

Greetings from the famous room N101 at ANU Computer Science where Prof Willy Zwaenepoel is speaking on "Software for Fast Storage Hardware".  He argues that modern solid state storage works differently to the old rotating disks database software was designed for. As a result the old software doesn't produce the expected improvement of much faster storage. 


The problem turns out to be that the CPU is holding up data access. For those brought up with mechanical storage it is a shock, as disks are thousands of times slower than the CPU, but modern storage is solid state, just like the CPU.

Database software is optimized for sequential writing to the disk, because that is faster than random access. Logs are also used to allow for a failure while the data is being written. This all takes CPU cycles to do, but is not needed for SSD. Instead data can be written immediately to storage. 

Professor Zwaenepoel's software is called KVell (Key Value and North American slang for happy and proud). This reminded me of what Canberra start-up Instaclustr, do with Apache Cassandra. There are many articles for Cassandra setting for SSDs.

This was an excellent seminar to be back on campus for. It challenged assumptions from my earliest training at the ABS decades ago. 


ps: ANU events follow Oxford Time: "Classes commence at five minutes past the published start time and conclude five minutes before the published end time."(Policy: Timetable, ANUParagraph 6, 2019).

 

Monday, July 8, 2019

Education and big data in Australia

Writing in EduResearch Matters, Buchanan and McPherson (2019), discuss banning of smart phones in public schools, companies collecting data about students, data collection in schools for educational purposes, and the monitoring of individual students performance using learning management systems. However, the authors have conflated related, but separate topics.

Bans on student mobile devices are intended to reduce student distraction. This has nothing to do with collection of data about students. I suggest it would be better to teach students, particularly older students, how to use mobile devices responsibly, than banning them. I am old enough to have been shown how to make an emergency phone call at school: is that still done?

Data collection via social media, and mobile devices by corporations is an issue, but not one exclusively for teachers. What is a school issue is the use of corporate educational sites which are “free”, but collect student data for resale. Teachers should not use Apps which infringe their students privacy.

Extensive standardized testing of students predates the Internet, but is facilitated by it, as in the example of online NAPLAN. What needs to be remembered is collecting data is not in itself useful. Also there has been extensive research on how such testing can be harmful.

The propensity of school systems to measure students and try to put their behavior (not just their academic knowledge), on some sort of scale is facilitated by a greater ability to collect data. But then again there should be a good reason and evidence, this actually works.
If the data is not being collected for a good educational reason, then I suggest teachers have a professional responsibility not to collect it.

Like many AARE articles, this one portrays teachers as powerless employees required to carry out the instructions of their employers. I suggest teachers need to assert their professional status, and decide what is in the interests of their clients (the students), as all professionals are ethically required to do. Where data collection is not educationally justified, or is harmful, teachers have an ethical obligation not to collect that data. Teachers need to put in place guidelines, and then lobby collectively to have them adopted by school systems.

References


Buchanan, R., & McPherson, A. (2019). Teachers and learners in a time of big data. Journal of Philosophy in Schools, 6(1). URL http://dx.doi.org/10.21913/JPS.v6i1.1566

Tuesday, May 1, 2018

Training the Big Data Trainer

Greetings from the Australian National University in Canberra, where I am attending "Powering up your 2018 (data skills) training". This is provided by Australian National Data Service (ANDS), National eResearch Collaboration Tools and Resources project (Nectar) and  Research Data Services (RDS). The idea is to help those who train researchers in using data repositories.

Data repositories are used not just by researchers, but also for teaching. The Atlas of Living Australia is also used by schools.

One of the problems with data repositories of data is motivating researchers to learn to use them and use them. The researcher tends to be focused on their PhD, or project and then getting the next grant. Making their data available for others to use may not be seen as helping with this (although researchers now get some credit for publishing data).

ps: The instructor perhaps was having a dig at me just now saying "You don't want people playing with computers and not paying attention", but obviously I am. ;-)

Wednesday, November 19, 2014

Australian Government Open Data Repository Focusing on Real Data

Greetings from  the "Big data, big opportunity" conference, at the Australian National University in Canberra, where Pia Waugh, Director of Coordination and Gov 2.0, Australian Government is explaining the open data strategy. She emphasized that real data for data.gov.au is needed, not electronic copies of government documents. A good example is that the budget data was provided as computer readable files, as well as human readable documents. Not all government data can be made public as it contains confidential information about people and organizations. I noticed Pia is the only one of the presenters talking about information access how actually has a computer in front of them.

ps: Recently the Australian Prime Minister said "Coal is good for humanity ...", perhaps we can get Malcolm Turnbull , Minister for Communications to say: "Data is good for humanity, data is good for prosperity, data is an essential part of our economic future, here in Australia, and right around the world". ;-)

Climate Governance

Greetings from  the "Big data, big opportunity" conference, at the Australian National University in Canberra, where Eliza Murray is speaking on "Could order and ambition emerge from the fragmented climate governance complex?". She is using network analysis and demonstrated this using Gephi software. The idea is to look at the connections between national and international institutions involved with climate change. She estimates there are about 1,000 institutions to research, which therefore requires automated analysis. She intends to mine the website of climate change organizations and look at the overlapping memberships. She expects to find that the organizations have become decentralized by not fragmented.

Eliza commented that she initially stated her research on Australian organizations, but found these were not indicative of the international situation. When interviewing people in other countries she is asked "Why are these weird things going on in Australia with climate change?

I suggested Eliza contact Paul Thomas and the people at CSIRO researching information retrieval.

Citizen Mapping in Indonesia

Greetings from  the "Big data, big opportunity" conference, at the Australian National University in Canberra, where Christina Griffin is speaking on "Open access spatial data for effective disaster risk reduction". She is emphasizing the use of open street map for dealing with disasters in developing nations. While crowd sourced mapping data has limitations, it is better than no data at all. She is studying vulnerability to disasters in central Java, Indonesia. It happens I helped with the deployment of Sahana open source software for the 2006 Yogjakarta Earthquake in Indonesia.

New Zealand Educational Entrepreneurs

Greetings from the Griffin Room, overlooking Lake Burley Griffin at the Australian National University in Canberra, where Steve Thomas is speaking on ‘Putting a Value On It’. The value that New Zealand educational entrepreneurs plan to create. He started by defining social entrepreneurship, in terms of innovation, revenue generation for improving welfare. He said there was not much research on this.

He is studying "Partnership Schools" (Kura Hourua)  (PSKH). These allow more flexibility in teaching, with different teaching hours and unregistered teachers. Examples included military style schools, farm based and Mauri and pacific inland culture orientated schools. Some schools plan to use Wraparound Services to address health and low socio-economic status.

Steve pointed out that some of these services are not new, but are being delivered in new ways. But it is too early to show this works.

I asked Steve about use of e-learning in NZ schools, given the NZ Education Department developed the Mahara e-portfolio software. He commented that several of the people interviewed had commented they were looking at IT use, but did not seem to be clear on how to do this.

Saturday, November 15, 2014

Big Data Policy Conference in Canberra

There are still some free tickets left for the "Big data, big opportunity" conference, at the Australian National University in Canberra, 19 November 2014. This features Pia Waugh, Director of Coordination and Gov 2.0, Australian Government. This is a PHD conference, where the program is mostly research students presenting their work in short snappy presentations (along with some keynotes by celebrities).
The Big Data, Big Opportunity conference will examine opportunities presented by effectively harnessing big data, and in particular will look at how open data (across the Academy, government and industry) can enhance research, shape policy development, and impact on innovation. The conference will provide a forum for PhD students, academics, policymakers, and industry to jointly discuss the implications and challenges of moving towards open access data.
The conference will run the following topic streams:
• Economics/economic policy
• Public policy and governance
• Environment, development and resource management

My picks for the day

Time Session
0845 Registration
0915 Welcome: Arjuna Mohottala, President, Crawford PhD Conference 2014 Organising Committee, Molonglo Theatre
0920 Keynote: John McMillan, Australian Information Commissioner, Molonglo Theatre
1010 Morning tea.
1030 Griffin Room: ‘Putting a Value On It’. The value that New Zealand educational entrepreneurs plan to create, Steve Thomas
1100 Lennox Room: Applying reinforcement learning to single and multi-agent economic problems, Neal Hughes
Discussant: Akshay Shanker
1130 Griffin Room: Where big data meets no data, Belinda Thompson
1200 Griffin Room: Facing our demons: Do mindfulness skills help people deal with failure at work? James Donald
1230 Lunch.
1300 Griffin Room: Giving rights to nature: A new institutional approach for overcoming social dilemmas? Julia Talbot-Jones
1330 Molonglo Theatre: Small states, big effects? Oil price shocks and economic growth in small island developing states, Alrick Campbell. Discussant: Arjuna Mohottala
1400 Griffin Room: Open access spatial data for effective disaster risk reduction, Christina Griffin
1430 Lennox Room: Could order and ambition emerge from the fragmented climate governance complex? Eliza Murray
1500 Afternoon tea.
1520 Panel session: The successes, challenges, and future potential of big data, Molonglo Theatre
Chair: Jenny Gordon, Principal Adviser Research Canberra, Productivity Commission
• Andy Heys, Software Architect, IBM Australia
• Greg Laughlin, Principal Policy Adviser, Australian National Data Service
• Duncan Stone, Senior Manager, Open Innovation, Price waterhouse Coopers
• Pia Waugh, Director of Coordination and Gov 2.0, Australian Government
1645 Wrap-up: Arjuna Mohottala, President, Crawford PhD Conference 2014 Organising Committee
1700 Close

Thursday, August 15, 2013

Privacy Preserving Data Integration Strategies

Professor Bradley Malin, Vanderbilt University (Nashville, USA), will speak on "Towards Practical Private Data Integration and Analysis" 4pm 26 August 2013, in the famous room N101 at the Australian National University in Canberra.

Towards Practical Private Data Integration and Analysis

Assoc Prof Bradley Malin (Vanderbilt University, Nashville)

DATE: 2013-08-26
TIME: 16:00:00 - 17:00:00
LOCATION: CSIT Seminar Room, N101

ABSTRACT:

Over the past decade, it has been repeatedly demonstrated that data devoid of explicit identifiers can be linked back to the identities of the individuals from which it was derived. This has made organizations increasingly apprehensive about sharing person-specific information. Yet, with the dawn of the big data age upon us, it is imperative that data sharing proliferate to ensure that researchers can validate published research findings, combine datasets to discover novel associations, and comply with open data initiatives. In this talk, I will review recent research on privacy preserving data integration strategies that are efficient, effective, and obscure personal identities in the process. This talk will further illustrate how such integration can enable biomedical association studies while obfuscating the identities of the corresponding participants.

BIO:

Bradley Malin, Ph.D., is the Vice Chair for Research and an Associate Professor of Biomedical Informatics in the School of Medicine at Vanderbilt University. He is also an Associate Professor of Computer Science in the School of Engineering and is Affiliated Faculty in the Center for Biomedical Ethics and Society. He is the founder and current director of the Health Information Privacy Laboratory (HIPLab), conducts technologies that enable privacy in the context of real world organizational, political, and health information architectures. Dr. Malin's research has been cited by the U.S. Federal Trade Commission and featured in popular media outlets, including Nature News, Scientific American, and Wired magazine. He has received several awards of distinction from the American and International Medical Informatics Associations and, in 2009, he was honored as a recipient of the Presidential Early Career Award for Scientists and Engineers (PECASE), the highest honor bestowed by the U.S. government on outstanding scientists and engineers beginning their independent careers. Dr. Malin completed his education at Carnegie Mellon University in Pittsburgh, PA, where he received a bachelor's in biological sciences, a master's in data mining and knowledge discovery, a master's in public policy and management, and a doctorate in computer science.