[{"content":" All blogs, on any topic\u0026hellip; ","date":"5 March 2025","externalUrl":null,"permalink":"/en/posts/","section":"Blogs","summary":" All blogs, on any topic… ","title":"Blogs","type":"posts"},{"content":"","date":"5 March 2025","externalUrl":null,"permalink":"/en/tags/ielts/","section":"Tags","summary":"","title":"Ielts","type":"tags"},{"content":" My 1st IELTS score: Overall 7 (Listening 8, Reading 8, Writing 6, Speaking 5.5) Summarized my strategies and exam techniques for IELTS preparation.\nFor basic information about the exam (such as the exam process, question composition, etc.), please find other materials on your own.\nThese are personal experiences — adapt what works for you. Section-Specific Strategies # Listening # Patterns: Each part follows fixed formats (academic/conversational, monologue/dialogue). Vocabulary: Focus on spelling, synonyms, and context comprehension. Expect unfamiliar words—learn to move past them. Note word limits for answers. Predict answer types (numbers/nouns/verbs/plurals) from questions. Highlight keywords to track your progress. If lost, use keywords to relocate. Use gaps to pre-scan later sections. Prioritize Part 3 keywords before Part 1 starts. Part 1: Drill common formats (postcodes, phone numbers, dates, currencies). Reading # Core principle: Match questions to the text. Vocabulary: Master high-frequency words (CET-4/6, IELTS core lists). Skip obscure terms—they’re rarely tested. Speed: Time yourself (10+15+25 minutes). Never linger on a question—all are worth 1 point. Techniques: Skim headings, questions, and paragraph openings/closings first. Parallel Reading: Tackle multiple question types in one pass if they follow text order. Combine skimming (find sections) and scanning (extract answers). Writing # Time: 15m for Task 1, 40m+5m planning for Task 2. The question types and structures of the essays are formulaic; the structure can be roughly adjusted as needed. Remember transition words and formulaic sentences, and practice your thought process Task1 tests descriptive ability, with the structure: Paraphrase + Overview + Detail1 + Detail2. Memorize the common words for each question type and practice how to identify key points Task2 tests persuasive ability, with the structure: Introduction + Body1 + Body2 + Conclusion. Learn the methods of presenting viewpoints + discussion + examples for argumentation Speaking # Goal: Think and speak naturally. Avoid over-rehearsed answers. Practice: Record yourself. Use fillers (well, you know) to buy time. Prioritize fluency over complex vocabulary. Part 1: Prepare short answers about daily life. Part 2: Reuse adaptable stories (people/events/objects). Part 3: Expand answers with brainstorming—examiners assess depth and coherence. Key Resources # Quickly review these and pick what suits you:\nIELTS Ready Premium: Official IELTS tool, free to use after registration, provides detailed courses and exercises. Great for familiarizing with question types, practicing speaking (one-on-one Q\u0026amp;A), and writing (sample essays for different scores). Highly recommended. IELTS Liz: Comprehensive website, offers numerous IELTS practice materials and tips, see the same-name YouTube Channel. IELTS-up: Another comprehensive website, particularly good for summarizing useful phrases for writing. Recommended for consciously using these phrases during writing practice. Teacher Sun\u0026rsquo;s IELTS Practice: Includes various IELTS-related knowledge, primarily focuses on hands-on training for listening and reading. Clear demonstration of strategies, keyword marking, and choices by experienced test takers. Shen Xiaoyi Series Books: Highly recommended books like \u0026ldquo;Breaking IELTS Writing in 10 Days\u0026rdquo; and \u0026ldquo;Breaking IELTS Speaking in 10 Days.\u0026rdquo; Buy physical copies for in-depth study. IELTS Bro: Excellent computer-based mock test software; eliminates the need to purchase paperback Cambridge IELTS books. The simulated experience is identical to the real test. Highly recommended for thorough practice of Cambridge IELTS materials. Maimemo Vocabulary Builder: Start memorizing high-frequency words early and avoid obscure ones. Use English-English explanations and example sentences to learn words in context. You can customize vocabulary lists from high-frequency words found online. AI IELTS Essay Correction: Other AI tools can also be used for low-cost, efficient checks for spelling, grammar, and idea errors. Immediate feedback available. Final Tips # Band scores Corresponding to Number of Mistakes: The IELTS official website and various educational websites have this information. Why are you focusing on this? Are you trying to control your score!? Just aim to do your best. Booking the Exam: Make sure to read all the announcements on the IELTS official website and book your exam date as early as possible. Set a deadline for yourself to avoid procrastination. Computer-based/Paper-based: If you are proficient enough with a standard keyboard, I recommend the computer-based test. You can use software to simulate and compare; there are many advantages, especially for writing where you can freely format and not worry about handwriting (I repeatedly adjusted paragraphs and sentence orders during the exam, and even wrote scattered sentences before combining them). Tutoring Classes: If you have sufficient self-motivation and information-searching skills, you can definitely skip these classes. Moreover, the cost of tutoring is often equivalent to several actual exams. Exam Day: Try to schedule the exam in the morning and the speaking test as early as possible in the afternoon. Cramming English at the last minute is not helpful; it\u0026rsquo;s better to take advantage of your morning energy to push through, avoiding excessive stress. Past Papers: I haven\u0026rsquo;t used them much, as I feel the Cambridge IELTS materials provide enough practice. However, for speaking, considering the changing question seasons, it might be useful to look at recent exam materials for their timeliness. Long-Term Learning # If you are pursuing the long-term cultivation of English ability or want to understand the essence and methods of language learning, 罗肖尼Shawney\u0026rsquo;s explanation (Comprehensible Input Hypothesis) is the answer:\nHow to Truly Master a Word Rethinking Language Learning: Listening/Speaking Rethinking Language Learning: Reading Overall, fluency in a foreign language stems from massive exposure to comprehensible listening and reading materials. It involves finding a series of materials that are moderately challenging (i+1, meaning slightly above your current level), interesting, and sufficient in quantity. Only through extensive input, akin to acquiring a sense of language as one does with their native tongue, can one naturally begin to speak and output the language.\n","date":"5 March 2025","externalUrl":null,"permalink":"/en/posts/ielts-preparation-guide/","section":"Blogs","summary":" My 1st IELTS score: Overall 7 (Listening 8, Reading 8, Writing 6, Speaking 5.5) Summarized my strategies and exam techniques for IELTS preparation.\n","title":"IELTS Preperation Guide","type":"posts"},{"content":"","date":"5 March 2025","externalUrl":null,"permalink":"/en/","section":"Nand Fun","summary":"","title":"Nand Fun","type":"page"},{"content":"","date":"5 March 2025","externalUrl":null,"permalink":"/en/tags/","section":"Tags","summary":"","title":"Tags","type":"tags"},{"content":"","date":"2025-03-05","externalUrl":null,"permalink":"/tags/%E9%9B%85%E6%80%9D/","section":"Tags","summary":"","title":"雅思","type":"tags"},{"content":"","date":"17 October 2024","externalUrl":null,"permalink":"/en/tags/configuration/","section":"Tags","summary":"","title":"Configuration","type":"tags"},{"content":"","date":"17 October 2024","externalUrl":null,"permalink":"/en/tags/macos/","section":"Tags","summary":"","title":"MacOS","type":"tags"},{"content":" Rational configuration and efficient functional software, creating a smooth and efficient workflow. The content will be explained in the order of the configuration. Some parts can be selectively configured according to individual needs.\nIt is recommended to read through the entire text first, and not to rush into the configuration. The in-depth configuration of some applications may depend on other applications, so it is absolutely crucial not to implement everything at once! Prerequisites # VPN (Optional) # The majority of computer-related applications are difficult to access in mainland China. If users in mainland China want to comfortably configure these applications, they must complete this step.\nCommand Line Tools (CLT) for Xcode # For macOS users, the system comes with bash, git, and curl pre-installed. In the command line, enter xcode-select --install to install CLT for Xcode.\nHomebrew # The most important application on macOS is Homebrew. Please learn the detailed usage methods of Homebrew\nHomebrew is the most widely used and best package manager on macOS. Basically, all \u0026ldquo;small\u0026rdquo; applications (except for large and robust ones like Chrome) are recommended to be downloaded and managed using the package manager.\nFor users in mainland China, it is recommended to install and configure Homebrew according to the Tsinghua University Open Source Software Mirror Site. It\u0026rsquo;s best to directly switch the source.\nIf mainland China users have a VPN, they can also directly install Homebrew using the terminal commands on the Homebrew official website.\nGitHub (Recommended) # Recommended: Register and configure a GitHub account with SSH protocol connection.\nBasic Applications # Zsh # First, use brew to install zsh and set it as the default shell. brew install zsh chsh -s $(which zsh) I use Oh My Zsh as a framework and package manager for managing the use of zsh, and install it according to the official website command line instructions. The Oh My Zsh plugin system is very rich, and you can refer to my .zshrc for recommended plugins to install or enable. Please make sure you understand the meaning of each line in the configuration file before adding it. Tools such as Starship can be used to enhance the prompt, and you can refer to my configuration file. Node.js # Recommend using fnm for managing multiple Node.js environments. Install it using brew and follow the configuration instructions on its official website.\nJava # Recommend using jenv for managing multiple Java environments. Install it using brew and configure it according to the official website.\nConda # Recommend installing Miniconda, which is more lightweight. You can install it according to the official documentation.\nRust # According to the official website, you can install Rust using the command line.\nGolang # According to the official website, you can install it using the command line.\nProductivity tools # Karabiner-Elements # A powerful and stable keyboard customizer\nRecommend downloading from the official website, it has powerful functions and provides a high degree of customization for both keyboard and mouse.\niTerm2 # iTerm2 is a terminal emulator, a replacement for Terminal and the successor to iTerm. It not only has a better appearance but also stronger functionality. Recommended configuration of hotkeys and themes.\nRecommended to use brew for installation, or you can also download and install it from the official website.\nNeoVim # hyperextensible Vim-based text editor, which I generally use for lightweight file editing on the command line.\nIt is recommended to use brew for installation. The in-depth configuration of NewVim is quite complex, but there are many pre-configured \u0026ldquo;distributions\u0026rdquo; available. Personally, I use LunarVim with some customization.\nLunarVim involves a large number of dependencies, please carefully follow its official documentation for installation. JetBrains # It is recommended to use the JetBrains Toolbox App for installing and managing JetBrains software products.\nVsCode # A lightweight and versatile editor\nRecommended to install from the official website.\nRaycast # Free and more powerful alternatives to Alfred\nRecommended to use brew for installation, or you can also download and install from the official website.\nThe plugin ecosystem is powerful and has a wide range of features. Be sure to browse the plugin marketplace thoroughly. Tuxera # The NTFS-formatted hard drive for Microsoft can be used on a Mac, but the personal use of that paid software has other alternative options.\nPeek # The Mac App Store is available for download, and macOS allows you to quickly preview files by pressing the spacebar. Peek has enhanced this feature, supporting a wider range of file formats.\niStat Menus # The ultimate system monitor\nThe new version interface is more visually appealing. The best detector for Mac, with customizable status bar and rich, detailed monitoring data.\nBetter365 # The official website is here, this company has a large number of practical Mac applications, the ones I use are as follows:\niShot: Screenshot, long screenshot, image pasting, annotation, color picking, screen recording tool FastZip: Excellent free compression and decompression tool Automatic input method switching: Can help you automatically switch input methods (Chinese and English) iBar: Hide menu bar icons Better And Better: Essential Mac tool for mouse, trackpad, and keyboard gestures draw.io # draw.io is free online diagram software for making flowcharts, process diagrams, org charts, UML, ER and network diagrams.\nYabai # A tiling window manager for macOS based on binary space partitioning\nThe main issue is that the window switching and management in macOS itself is too slow and cumbersome. I really like its multi-space design, but the switching process is too troublesome.\nyabai has a complete Wiki, and you can also refer to my yabai configuration files.\nSketchyBar # A highly customizable macOS status bar replacement\nHere is the translation from Chinese to English:\nGitHub Link is available here, which can be used in conjunction with Yabai, with rich functionality. You can refer to my configuration file.\n","date":"17 October 2024","externalUrl":null,"permalink":"/en/posts/ultimate-macos-setup/","section":"Blogs","summary":" Rational configuration and efficient functional software, creating a smooth and efficient workflow. The content will be explained in the order of the configuration. Some parts can be selectively configured according to individual needs.\n","title":"Ultimate macOS Setup","type":"posts"},{"content":"","date":"2024-10-17","externalUrl":null,"permalink":"/tags/%E9%85%8D%E7%BD%AE/","section":"Tags","summary":"","title":"配置","type":"tags"},{"content":"I always try to squeeze in time for work and picking up new skills. Most of these pet projects of mine don’t end up seeing the light of day, to be honest. They are, however, great opportunities to try something in the real world and learn from it.\nProjects # Logo Link Description Shuohua A lightweight macOS voice input tool with real-time transcription, optional refinement, and automatic text insertion. Owl Split wireless keyboard for comfort and flexibility. Vision Guard An Intelligent Monitoring and Security System. Subway Route Finder A fast and convenient Chongqing subway route query system. Budget Buddy Personal finance manager: import WeChat/Alipay bills, analyze, generate reports. Moon Phase Tracker Explore moon phases for any selected date. Word Wise File manager organizes text files by keyword-based topics. Experience # Company Link Role Dates Location 4Paradigm Software Developer (Internship) 2026.04 - 2026.08 Shenzhen, Guangdong, China Digiwin Backend Developer (Internship) 2024.07 - 2024.10 Nanjing, Jiangsu, China iFlytek Software Developer (Internship) 2024.03 - 2024.05 Chongqing, China Education # School Link Degree Date CUHK-Shenzhen MSc, Computer Science 2025 - 2027 Southwest University BEng, Computer Science 2021 - 2025 Wenzhou High School High School Education 2018 - 2021 Wenzhou Foreign Language School Middle School Education 2015 - 2018 ","date":"25 September 2024","externalUrl":null,"permalink":"/en/about/","section":"Nand Fun","summary":"I always try to squeeze in time for work and picking up new skills. Most of these pet projects of mine don’t end up seeing the light of day, to be honest. They are, however, great opportunities to try something in the real world and learn from it.\n","title":"About","type":"page"},{"content":"","date":"15 August 2024","externalUrl":null,"permalink":"/en/tags/keyboard/","section":"Tags","summary":"","title":"Keyboard","type":"tags"},{"content":" \u003c?xml version=\"1.0\" encoding=\"UTF-8\" ?\u003e \u003c!DOCTYPE svg PUBLIC \"-//W3C//DTD SVG 1.1//EN\" \"http://www.w3.org/Graphics/SVG/1.1/DTD/svg11.dtd\"\u003e Please note that if you encounter any difficulties during the installation process, please fully utilize your DIY abilities and use various feasible parts. Switch Socket # Please select the corresponding keyboard switch socket according to the selected PCB.\nMX switch and Gateron low-profile switch for normal version, MX switch and Kailh low-profile switch for choc version\nRGB # A single side needs 23 key lights and 6 bottom lights in series, so please install all or none at all. For the convenience of selection and manual soldering, we only use SK6812MINI-E. The position for installing the GND pin has been marked on the PCB. When installing, just pay attention to have the LED facing downwards and the bottom light facing upwards.\n\u003c?xml version=\"1.0\" encoding=\"UTF-8\" ?\u003e \u003c!DOCTYPE svg PUBLIC \"-//W3C//DTD SVG 1.1//EN\" \"http://www.w3.org/Graphics/SVG/1.1/DTD/svg11.dtd\"\u003e They are in opposite directions Diode # Each keyboard switch corresponds to a diode. You can freely choose between surface-mount diodes or through-hole diodes, and even install them on the front or back of the PCB (note that installing on the back may conflict with the case).\n\u003c?xml version=\"1.0\" encoding=\"UTF-8\" ?\u003e \u003c!DOCTYPE svg PUBLIC \"-//W3C//DTD SVG 1.1//EN\" \"http://www.w3.org/Graphics/SVG/1.1/DTD/svg11.dtd\"\u003e Please make sure that the direction of the diode is correct Switch and Battery # The reset and boot switches, as well as the power socket, are very easy to solder. It is not necessary to elaborate here. It is worth noting that the protruding pins on the back of the straight plug switch should be polished and trimmed to avoid affecting the installation of the housing.\n\u003c?xml version=\"1.0\" encoding=\"UTF-8\" ?\u003e \u003c!DOCTYPE svg PUBLIC \"-//W3C//DTD SVG 1.1//EN\" \"http://www.w3.org/Graphics/SVG/1.1/DTD/svg11.dtd\"\u003e Please be careful to pay attention to the battery wire sequence Chip and Oled # Since the PCB is double-sided, you can install the chip on either side, whether it is directly soldered or connected with a socket. The position where the chip pins should be installed is marked on each side, be careful not to solder incorrectly.\n\u003c?xml version=\"1.0\" encoding=\"UTF-8\" ?\u003e \u003c!DOCTYPE svg PUBLIC \"-//W3C//DTD SVG 1.1//EN\" \"http://www.w3.org/Graphics/SVG/1.1/DTD/svg11.dtd\"\u003e If you want to use 3D printed shell, be sure to follow the steps below! It is recommended to use the low-profile slot to install the chip on the front of the keyboard. Pay attention to installing the slot on the column marked on the PCB. When installing the chip, make sure that its back side (the flat side without components) faces up. Finally, when installing the OLED screen, please make sure they are installed tightly.\nCase # 3D Printed # First, install five hot-melt screws into the 3D printed case, then place the completed PCB inside, tighten the five screws, and finally install the back panel and front OLED acrylic.\nTransparent Explorer # This case does not have much to pay attention to, only using m2 screws and two-stage screws for connection.\nKey # The final step is to install the keyboard switches and keycaps, just align them, very simple, please feel free to use your creativity to mix and match.\nFirmware # This keyboard uses ZMK as the firmware. Please fork my zmk-config repository and the zmk official documentation. It is recommended to use GitHub Action and keymap-editor for web-based visual firmware building.\nThe method of flashing firmware is very simple. Just press the boot button twice quickly to enter boot mode, and the keyboard controller will appear on the computer like a USB flash drive. Drag and drop the firmware to update it.\n","date":"15 August 2024","externalUrl":null,"permalink":"/en/posts/owl-keyboard-assembly/","section":"Blogs","summary":" \u003c?xml version=\"1.0\" encoding=\"UTF-8\" ?\u003e \u003c!DOCTYPE svg PUBLIC \"-//W3C//DTD SVG 1.1//EN\" \"http://www.w3.org/Graphics/SVG/1.1/DTD/svg11.dtd\"\u003e Please note that if you encounter any difficulties during the installation process, please fully utilize your DIY abilities and use various feasible parts. Switch Socket # Please select the corresponding keyboard switch socket according to the selected PCB.\n","title":"Owl Keyboard: Assembly","type":"posts"},{"content":"","date":"15 August 2024","externalUrl":null,"permalink":"/en/series/owl-split-keyboard/","section":"Series","summary":"","title":"Owl Split Keyboard","type":"series"},{"content":"","date":"2024-08-15","externalUrl":null,"permalink":"/series/owl-%E5%88%86%E4%BD%93%E9%94%AE%E7%9B%98/","section":"Series","summary":"","title":"Owl 分体键盘","type":"series"},{"content":"","date":"15 August 2024","externalUrl":null,"permalink":"/en/series/","section":"Series","summary":"","title":"Series","type":"series"},{"content":"","date":"15 August 2024","externalUrl":null,"permalink":"/en/tags/zmk/","section":"Tags","summary":"","title":"ZMK","type":"tags"},{"content":"","date":"2024-08-15","externalUrl":null,"permalink":"/tags/%E9%94%AE%E7%9B%98/","section":"Tags","summary":"","title":"键盘","type":"tags"},{"content":" \u003c?xml version=\"1.0\" encoding=\"UTF-8\" ?\u003e \u003c!DOCTYPE svg PUBLIC \"-//W3C//DTD SVG 1.1//EN\" \"http://www.w3.org/Graphics/SVG/1.1/DTD/svg11.dtd\"\u003e Please read the guide carefully. If possible, please also refer to other split keyboard construction guides as reference. PCB # Name Count Remarks PCB 2 1.6mm if use 3D printed case ProMicro 2 nice!nano Diodes 46 SOD-123 or DO-35 Reset switch 2 3*6*5 Battery switch 2 MSK-1102-1.5H Battery socket 2 PH2.0 2P Battery 2 3.7V lithium (max size 40*50*70) PCB sockets 46 Compatible with MX and Gateron low profile, and Kailh choc if using choc version PCB Key switches 46 Only compatible with MX style Keycaps 46 1u cap Previous Next Optional # RGB and Oled # Name Count Remarks OLED 2 0.91 inches SK6812MINI-E 58 Install all lights as they are connected in series Previous Next Case # General # Name Count Remarks Spacer M2 12mm 4 For OLED cover Screw M2 24 At least 3mm Rubber feet ANY For anti-slip purposes Transparent Explorer # Name Count Remarks Top plate 2 1.5mm-4mm Bottom plate 2 At least 1.5mm OLED cover 2 At least 1mm Spacer M2 7mm 10 Between pcb and bottom plate 3D Printed # Name Count Remarks Spacer M2 7+4mm 10 Between pcb and bottom plate Hot melt screw M2*3*3.5 10 Used for integrating 3D printed case 3D-printed case\n2 Please select printing materials freely OLED cover 2 2mm Bottom plate 2 2mm ","date":"14 August 2024","externalUrl":null,"permalink":"/en/posts/owl-keyboard-materials/","section":"Blogs","summary":" \u003c?xml version=\"1.0\" encoding=\"UTF-8\" ?\u003e \u003c!DOCTYPE svg PUBLIC \"-//W3C//DTD SVG 1.1//EN\" \"http://www.w3.org/Graphics/SVG/1.1/DTD/svg11.dtd\"\u003e Please read the guide carefully. If possible, please also refer to other split keyboard construction guides as reference. PCB # Name Count Remarks PCB 2 1.6mm if use 3D printed case ProMicro 2 nice!nano Diodes 46 SOD-123 or DO-35 Reset switch 2 3*6*5 Battery switch 2 MSK-1102-1.5H Battery socket 2 PH2.0 2P Battery 2 3.7V lithium (max size 40*50*70) PCB sockets 46 Compatible with MX and Gateron low profile, and Kailh choc if using choc version PCB Key switches 46 Only compatible with MX style Keycaps 46 1u cap Previous Next Optional # RGB and Oled # Name Count Remarks OLED 2 0.91 inches SK6812MINI-E 58 Install all lights as they are connected in series Previous Next Case # General # Name Count Remarks Spacer M2 12mm 4 For OLED cover Screw M2 24 At least 3mm Rubber feet ANY For anti-slip purposes ","title":"Owl Keyboard: Materials","type":"posts"},{"content":" Owl (Orthogonal Wireless Layout) \u003c?xml version=\"1.0\" encoding=\"UTF-8\" ?\u003e \u003c!DOCTYPE svg PUBLIC \"-//W3C//DTD SVG 1.1//EN\" \"http://www.w3.org/Graphics/SVG/1.1/DTD/svg11.dtd\"\u003e Combines the ergonomic freedom of wireless connectivity with a split keyboard Design influenced by Lily58, Corne, Sofle and Torn keyboards\nFeature # Owl uses Pro Micro Interconnect small controller board, which means easy for Maintenance and replacemen. And supports a variety of boards including nice!nano and nrfmicro (I have not tried yet, but theoretically feasible). Low latency wireless, supported by ZMK. Gorgeous RGB lighting effects. Two multifunctional OLED screens that can display keyboard battery level and connection status. Split keyboard conforms to ergonomics. A total of two PCBs compatible with Cherry MX switches, Kailh choc Low Profile switches, and Gateron Low Profile switches. Double-sided PCB design, both sides are universal. DIY # Please make sure you can get the ability to build the keyboard:\nSoldering circuit components Circuit board production 3D printing Using GitHub action to build firmware Consulting documents carefully \u003c?xml version=\"1.0\" encoding=\"UTF-8\" ?\u003e \u003c!DOCTYPE svg PUBLIC \"-//W3C//DTD SVG 1.1//EN\" \"http://www.w3.org/Graphics/SVG/1.1/DTD/svg11.dtd\"\u003e Continue reading the Owl Split Keyboard series to make a keyboard License # The Owl Keyboard is licensed under Creative Commons Attribution-NonCommercial 4.0 International License.\nThis PCB design may be freely reproduced, modified, and manufactured for personal use only. If you would like to use this design commercially please contact me at hza2002@foxmail.com to request permission.\n","date":"13 August 2024","externalUrl":null,"permalink":"/en/posts/owl-keyboard-overview/","section":"Blogs","summary":" Owl (Orthogonal Wireless Layout) \u003c?xml version=\"1.0\" encoding=\"UTF-8\" ?\u003e \u003c!DOCTYPE svg PUBLIC \"-//W3C//DTD SVG 1.1//EN\" \"http://www.w3.org/Graphics/SVG/1.1/DTD/svg11.dtd\"\u003e Combines the ergonomic freedom of wireless connectivity with a split keyboard Design influenced by Lily58, Corne, Sofle and Torn keyboards\n","title":"Owl Keyboard: Overview","type":"posts"},{"content":"","date":"10 January 2023","externalUrl":null,"permalink":"/en/tags/convolution/","section":"Tags","summary":"","title":"Convolution","type":"tags"},{"content":"","date":"10 January 2023","externalUrl":null,"permalink":"/en/tags/image-filtering/","section":"Tags","summary":"","title":"Image Filtering","type":"tags"},{"content":"","date":"10 January 2023","externalUrl":null,"permalink":"/en/tags/image-processing/","section":"Tags","summary":"","title":"Image Processing","type":"tags"},{"content":"","date":"10 January 2023","externalUrl":null,"permalink":"/en/tags/linear-algebra/","section":"Tags","summary":"","title":"Linear Algebra","type":"tags"},{"content":" Image is composed of individual pixels, and the essence of image processing is to process these pixels. The correlation between pixels is important information, and it cannot be completely separated, which is also the starting point of many image algorithms. Convolution is a common operation in mathematics. In image processing, it changes the details of the image by performing operations between the image and the convolution kernel. Convolution or filtering operations on images are widely used in various scenarios, such as various filters and convolutional neural networks. This article will explore the knowledge of linear algebra applied in image processing convolution through image filtering. Introduction to Image Processing # Image processing refers to the analysis, processing, and processing of images to meet visual, psychological, or other requirements. Image processing is an application of signal processing in the field of images. Currently, most images are stored in digital form, so in many cases, image processing refers to digital image processing. In addition, processing methods based on optical theory still occupy an important position.\nImages in computers are composed of a large number of seemingly continuous pixels, and mathematically, each pixel of an image can correspond to each element of a matrix in linear algebra, so images can be represented using matrices. The types of images vary, and the dimensions of the matrices will change: a grayscale image can be represented by a two-dimensional matrix with element values ranging from 0 to 255, where the element value corresponds to the brightness of the pixel (0 corresponds to black, 255 corresponds to white); a color image (RGB image) can be represented by a three-dimensional matrix, with the red (R), green (G), and blue (B) components represented by three separate matrices, and the combination of the three matrices forms the three-dimensional matrix. It can be said that an image is equivalent to a matrix, so the results of matrix theory in linear algebra can be applied to image processing.\nSimple Geometric Transformations # Treating image information as a matrix for processing, we can utilize the knowledge of matrix theory in linear algebra to perform simple geometric transformations on images. The following introduces several common principles of image geometric transformations, where \\(x\u0026rsquo;\\), \\(y\u0026rsquo;\\) are the pixel coordinates of the transformed image, \\(x_0\\), \\(y_0\\) are the offsets in each direction, and \\(x\\), \\(y\\) are the pixel coordinates of the original image.\nTranslation # $$ \\begin{bmatrix} x\u0026rsquo; \\cr y\u0026rsquo; \\cr 1 \\end{bmatrix} = \\begin{bmatrix} 1\u0026amp;0\u0026amp;x_0\\cr 0\u0026amp;1\u0026amp;y_0\\cr 0\u0026amp;0\u0026amp;1 \\end{bmatrix} \\begin{bmatrix} x \\cr y \\cr 1 \\end{bmatrix} $$\nRotation # $$ \\begin{bmatrix} x\u0026rsquo; \\cr y\u0026rsquo; \\cr 1 \\end{bmatrix} = \\begin{bmatrix} cos\\theta\u0026amp;-sin\\theta\u0026amp;0 \\cr sin\\theta\u0026amp;cos\\theta\u0026amp;0 \\cr 0\u0026amp;0\u0026amp;1 \\end{bmatrix} \\begin{bmatrix} x \\cr y \\cr 1 \\end{bmatrix} $$\nRotation may not be able to map perfectly to each new pixel, so post-processing of the rotated image is required for repair, but usually due to the lack of missing points, the RGB values of the previous pixel can be used to achieve this.\nScale # $$ \\begin{bmatrix} x\u0026rsquo; \\cr y\u0026rsquo; \\cr 1 \\end{bmatrix} = \\begin{bmatrix} c\u0026amp;0\u0026amp;0 \\cr 0\u0026amp;d\u0026amp;0 \\cr 0\u0026amp;0\u0026amp;1 \\end{bmatrix} \\begin{bmatrix} x \\cr y \\cr 1 \\end{bmatrix} $$\nEvidently, after magnification, there will be blank spaces without image information, so these spaces need to be supplemented. We can use bilinear interpolation to perform interpolation and supplement the output image.\nBilinear Interpolation Locate the nearest pixel points with valid RGB information in the four directions of the upper left, upper right, lower left, and lower right, and calculate their distances \\(d_{ul}\\), \\(d_{ur}\\), \\(d_{dl}\\), \\(d_{dr}\\) to the pixel point to be interpolated. Based on their inverse relationship, assign corresponding weights to the corresponding points.\n$$p_{ul}:p_{ur}:p_{dl}:p_{dr}= {1\\over d_{ul}}:{1\\over d_{ur}}:{1\\over d_{dl}}:{1\\over d_{dr}}$$\nMap the weights to \\([0,1]\\), then perform weighted assignment for each channel of the points to be interpolated:\n$$\\begin{bmatrix} R\u0026rsquo; \\cr G\u0026rsquo; \\cr B' \\end{bmatrix}=p_{ul}\\begin{bmatrix} R_{ul} \\cr G_{ul} \\cr B_{ul} \\end{bmatrix}+p_{ur}\\begin{bmatrix} R_{ur} \\cr G_{ur} \\cr B_{ur} \\end{bmatrix}+p_{dl}\\begin{bmatrix} R_{dl} \\cr G_{dl} \\cr B_{dl} \\end{bmatrix}+p_{dr}\\begin{bmatrix} R_{dr} \\cr G_{dr} \\cr B_{dr} \\end{bmatrix}$$\nThe interpolation of the entire image can be completed by performing the above operations on all the gaps.\nShear # shear on axis-x \\(\\begin{bmatrix} x\u0026rsquo; \\cr y\u0026rsquo; \\cr 1 \\end{bmatrix}=\\begin{bmatrix} 1\u0026amp;d_x\u0026amp;0 \\cr 0\u0026amp;1\u0026amp;0 \\cr 0\u0026amp;0\u0026amp;1 \\end{bmatrix}\\begin{bmatrix} x \\cr y \\cr 1 \\end{bmatrix}\\)\nshear on axis-y \\(\\begin{bmatrix} x\u0026rsquo; \\cr y\u0026rsquo; \\cr 1 \\end{bmatrix}=\\begin{bmatrix} 1\u0026amp;0\u0026amp;0 \\cr d_y\u0026amp;1\u0026amp;0 \\cr 0\u0026amp;0\u0026amp;1 \\end{bmatrix}\\begin{bmatrix} x \\cr y \\cr 1 \\end{bmatrix}\\)\nMirror # $$ \\begin{bmatrix} x\u0026rsquo; \\cr y\u0026rsquo; \\cr 1 \\end{bmatrix}= \\begin{bmatrix} s_x\u0026amp;0\u0026amp;0 \\cr 0\u0026amp;s_y\u0026amp;1 \\cr 0\u0026amp;0\u0026amp;1 \\end{bmatrix} \\begin{bmatrix} x \\cr y \\cr 1 \\end{bmatrix} $$\nConvolution # Principle # Convolution is the result of multiplying two variables within a certain range and then summing them. The term \u0026ldquo;convolution\u0026rdquo; first appeared in signal and linear system theory, where the focus is on the changes that a signal undergoes when passing through a linear system. Due to the common occurrence where the output of a signal at a previous moment affects the output at the current moment, the unit response of the system and the system input are generally used to find the convolution, in order to obtain the output signal of the system.\nOf course, this requires the system to be linear and time-invariant). If the convolution variables are the sequences \\(x(n)\\) and \\(h(n)\\), the result of the convolution is:\n$$ y(n)=\\sum_{i=-\\infty}^{\\infty} x(i) h(n-i)=x(n) * h(n) $$\nImage processing convolution # In image processing, convolution can be used for image filtering, feature extraction, and other processing. Mathematically, the convolution operation can be described using linear algebra methods, which is also a type of affine transformation: a linear transformation of the input vector.\nConvolution satisfies the definition of a linear function. If \\(f(𝑥)\\) satisfies the following two conditions, then \\(f(𝑥)\\) is said to be a linear function:\n$$ \\begin{aligned} 𝑓(𝑥1+𝑥2) \u0026amp;= 𝑓(𝑥1)+𝑓(𝑥2) \\cr 𝑓(𝑎𝑥) \u0026amp;= 𝑎𝑓(𝑥) \\end{aligned} $$\nDue to the fact that a digital image is a two-dimensional discrete signal, performing convolution on a digital image actually involves using the convolution kernel to slide across the image, multiplying the pixel values on the image with the corresponding values in the convolution kernel, and then summing all the multiplication results as the pixel value corresponding to the center of the convolution kernel. After sliding across all the pixels, the value of each pixel in the new image is the sum of the products of the pixel values in the original image and the weights of the convolution kernel.\nGiven an input image \\(f(x,y)\\) and a convolution kernel \\(g(x,y)\\), the result of the convolution operation is a new image \\(h(x,y)\\), where\n$$ h(x,y) = ∑∑ f(u,v) * g(x-u, y-v) $$\nThe convolution kernel \\(g(x,y)\\) is a small matrix used to perform convolution operations on the input image, and the output image \\(h(x,y)\\) is a new image obtained by the convolution of \\(g(x,y)\\) and \\(f(x,y)\\). Let \\(H\\) be the result of the convolution operation, which is a new matrix, \\(F\\) be the matrix representation of the input image, and \\(G\\) be the matrix representation of the convolution kernel, then the convolution operation can be represented as: \\(H = G * F\\).\nImage Filtering # The purpose of image filtering is to filter out noise in the image while preserving the image features as much as possible. It is an indispensable operation in image preprocessing, and its filtering effect directly affects the performance of subsequent image recognition, analysis, and other algorithms.\nThe essence of image filtering is to apply a convolution kernel to the image. Through different convolution kernels, different image processing effects can be achieved, such as edge enhancement, noise removal, and feature extraction.\nHere is the translation from Chinese to English:\nSome summary of convolutional kernels:\nConvolutional kernels are often matrices with odd numbers of rows and columns, which helps to better locate the center. The sum of the kernel elements reflects the brightness of the output. If the sum is 1, the brightness of the convolved image is similar to the original image; if the sum is 0, the convolved image is mostly black, and the brighter parts often represent certain features extracted from the image. High-pass filters (HPF) only allow the high-frequency parts (i.e., the parts with drastic changes) of the image to pass through, often used for image sharpening and enhancing the edges of objects in the image. Low-pass filters (LPF) only allow the low-frequency parts (i.e., the parts with gentle changes) of the image to pass through, often used for image blurring/smoothing and noise removal. Below are some examples of image filtering methods with their corresponding effect illustrations for reference. Image mean filtering # (x\u0026rsquo;, y\u0026rsquo;) are the pixel coordinates of the translated image, (x_0, y_0) are the offsets in each direction, and (x, y) are the pixel coordinates of the original image.\nThe main application of the mean filter is to remove irrelevant details in the image, that is, the pixel regions that are smaller than the size of the filter mask. This is achieved by taking the average of each pixel and its surrounding pixels within a mask (e.g., a 3×3 mask) and then reassigning the value.\n$$ Pixel\u0026rsquo;={1\\over9}\\times \\begin{bmatrix} 1\u0026amp;1\u0026amp;1 \\end{bmatrix} \\begin{bmatrix} 1\\cdot P_{ul} \u0026amp; 1\\cdot P_{up} \u0026amp; 1\\cdot P_{ur} \\cr 1\\cdot P_{l} \u0026amp; 1\\cdot P \u0026amp; 1\\cdot P_{r} \\cr 1\\cdot P_{dl} \u0026amp; 1\\cdot P_{down} \u0026amp; 1\\cdot P_{dr} \\end{bmatrix} \\begin{bmatrix} 1 \\cr 1 \\cr 1 \\end{bmatrix} $$\nImage Mean Filtering Example Laplace Image Enhancement # Principle: The Laplace operator can approximate the second-order derivative of an image, which has rotational invariance, meaning it can detect edges in all directions.\nThe Laplace operator\n$$ \\nabla^2f={\\partial^2f\\over\\partial x^2}+{\\partial^2f\\over\\partial y^2} $$\nIt includes:\n$$ \\begin{aligned} {\\partial^2f\\over\\partial x^2}=f(x+1,y)+f(x-1,y)-2f(x,y) \\cr {\\partial^2f\\over\\partial y^2}=f(x,y+1)+f(x,y-1)-2f(x,y) \\end{aligned} $$\nAfter incorporating the diagonal direction into the discussion\n$$ \\begin{aligned} {\\nabla^2f}=[f(x-1,y-1)+f(x,y-1)+f(x-1,y)+f(x+1,y)+\\cr f(x-1,y+1)+f(x,y+1)+f(x+1, y+1)]-8f(x,y) \\end{aligned} $$\nDue to the sudden change in pixel information values at the edges, i.e., the first-order derivative has a local maximum, this feature can be used to further calculate the second-order derivative equal to 0, and the above formula is the result of discretizing the second-order derivative. By using the Laplacian operator, the parts of the image where the information value (a certain RGB channel value) changes abruptly can be extracted, and this part is the edge of the image. Finally, by overlaying this with the original image, an edge-enhanced (sharpened) image is obtained. (Where \\(Pixel\u0026rsquo;\\) is the pixel coordinate of the transformed image)\nFiltering: $$ Pixel\u0026rsquo;={1\\over9} \\times \\begin{bmatrix} 1\u0026amp;1\u0026amp;1 \\end{bmatrix} \\begin{bmatrix} 1\\cdot P_{ul} \u0026amp; 1\\cdot P_{up} \u0026amp; 1\\cdot P_{ur} \\cr 1\\cdot P_{l} \u0026amp; -8\\cdot P \u0026amp; 1\\cdot P_{r} \\cr 1\\cdot P_{dl} \u0026amp; 1\\cdot P_{down} \u0026amp; 1\\cdot P_{dr} \\end{bmatrix} \\begin{bmatrix} 1 \\cr 1 \\cr 1 \\end{bmatrix} $$\nSuperimpose: $$ Pixel\u0026rsquo;\u0026rsquo;=Pixel - Pixel' $$\nLaplacian Image Enhancement Example ","date":"10 January 2023","externalUrl":null,"permalink":"/en/posts/linear-algebra-in-imaging/","section":"Blogs","summary":" Image is composed of individual pixels, and the essence of image processing is to process these pixels. The correlation between pixels is important information, and it cannot be completely separated, which is also the starting point of many image algorithms. Convolution is a common operation in mathematics. In image processing, it changes the details of the image by performing operations between the image and the convolution kernel. Convolution or filtering operations on images are widely used in various scenarios, such as various filters and convolutional neural networks. This article will explore the knowledge of linear algebra applied in image processing convolution through image filtering. Introduction to Image Processing # Image processing refers to the analysis, processing, and processing of images to meet visual, psychological, or other requirements. Image processing is an application of signal processing in the field of images. Currently, most images are stored in digital form, so in many cases, image processing refers to digital image processing. In addition, processing methods based on optical theory still occupy an important position.\n","title":"Linear Algebra in Imaging","type":"posts"},{"content":"","date":"2023-01-10","externalUrl":null,"permalink":"/tags/%E5%8D%B7%E7%A7%AF/","section":"Tags","summary":"","title":"卷积","type":"tags"},{"content":"","date":"2023-01-10","externalUrl":null,"permalink":"/tags/%E5%9B%BE%E5%83%8F%E5%A4%84%E7%90%86/","section":"Tags","summary":"","title":"图像处理","type":"tags"},{"content":"","date":"2023-01-10","externalUrl":null,"permalink":"/tags/%E5%9B%BE%E5%83%8F%E6%BB%A4%E6%B3%A2/","section":"Tags","summary":"","title":"图像滤波","type":"tags"},{"content":"","date":"2023-01-10","externalUrl":null,"permalink":"/tags/%E7%BA%BF%E6%80%A7%E4%BB%A3%E6%95%B0/","section":"Tags","summary":"","title":"线性代数","type":"tags"},{"content":" The 100 prisoners problem is a counterintuitive problem. It describes a seemingly impossible event: 100 prisoners have a chance to do the same thing, and only when all the prisoners do it can they survive. In fact, there is a reasonable implementation method for this problem to increase its probability by nearly 30 orders of magnitude. Introduction # The 100 prisoners problem is a mathematical problem in probability theory and combinatorics. In this problem, 100 numbered prisoners must find their own numbers in one of 100 drawers in order to survive. The rules state that each prisoner may open only 50 drawers and cannot communicate with other prisoners. Danish computer scientist Peter Bro Miltersen first proposed the problem in 2003. As an upgraded version of this problem, there will be \\(2N\\) prisoners. Their corresponding \\(2N\\) number cards are shuffled and placed in \\(2N\\) drawers. Each prisoner must open at most half of all the drawers and find the corresponding number card for his own number. All prisoners will enter the room separately. The failure of any one prisoner will result in the failure of the entire challenge. What is the maximum probability that the prisoners winning?\nTheoretical Analysis # Open the drawer randomly # Suppose All Prisoners Randomly Open \\(50\\) drawers In \\(100\\) Boxes.\nSet a random variable \\(X_i\\) as follows.\n$$ X_i=\\begin{cases} 1\\quad \\text{There is number i note in 50 drawers.}\\cr 0\\quad \\text{There is no number i note in 50 drawers.} \\end{cases} $$\nBecause the drawer where the note with its own number is uniformly distributed, the probability of pulling the box with its own numbered note is \\(P(X_i=1)=\\frac{50}{100}=\\frac{1}{2}\\) .\nAnd because the incidents in which all prisoners open the drawer are independent, the probability that all prisoners will get a note with their own number can be calculated as follows.\n$$ \\begin{aligned} P(X_1=1,X_2=1,\\cdots,X_{100}=1)=\u0026amp;P(X_1=1) \\cdots P(X_{100}=1)\\cr =\u0026amp;0.5^{100}\\cr \\approx \u0026amp;7.89 \\times 10^{-31} \\end{aligned} $$\nThis probability is obviously too low, even lower than selecting a lucky man out of \\(10^{30}\\) prisoners.\nOpen the drawer in a way # Theorem 1: If the number of a box is treated as a value and the number of a note is treated as a pointer, then there are and only are several cycles in the room.\nProof Because the numbers of the boxes and notes range from 1 to 100, each box has and only has one slip corresponding to its number and each slip is placed in a box, there are no slips without boxes or boxes without slips, that is, there are only cycles in the room. Theorem 2: If there is no limit on the number of steps, each person can find the note with their own number by the above method, that is, the note with their own number must be in the cycle they walk.\nProof From Theorem 1, since they start with the box with their own number, there must be a note with their own number pointing to the box with their own number in this cycle, that is, the note with their own number can be found if there is no limit on the number of steps. Now we can transform the problem into given that the prisoners act according to the above method, find the probability \\(P\\) that all prisoners succuss. From Theorem 2, we can further abstract the problem as: find the probability \\(P\\) that the length of all cycles in the \\(100\\) nodes with randomly assigned numbers from \\(1\\) to \\(100\\) is not more than \\(50\\).\nGeneral problem # Through the theoretical analysis of the special problem, that is, the problem of \\(100\\) prisoners, we further analyze the general solution to the problem, that is, the best strategy for \\(2N\\) prisoners.\nWhen there are \\(2N\\) prisoners, through Open the drawer in a way analysis, we can give the following optimal strategy. \\(2N\\) prisoners label \\(2N\\) drawers with random numbers. When each prisoner enters the room, they choose the drawer with the same number as their label, take out the note, and continue opening the drawer with the same number as the note, and so on.\nIf the number on the slip in drawer \\(k\\) is \\(s\\), then this strategy can be represented as \\(f(s)=k\\) , where \\(f\\) is a map from the drawer number to the number number, that is, from the current drawer to the next drawer. And this map has the following properties:\n$$ \\exists i, 1 \\leq i \\leq 2N, \\underbrace{f \\circ \\cdots \\circ f}_{i \\text{ times}}(s) = s $$\nCurrently, the condition for the prisoner\u0026rsquo;s success is:\n$$ \\forall s, 1 \\leq s \\leq 2N, \\exists i, 1 \\leq i \\leq N, \\underbrace{f \\circ \\cdots \\circ f}_{i \\text{ times}}(s) = s $$\nFor any number, it can be mapped up to itself by applying the function \\(i\\) times.\nTherefore, the sample space for this strategy is the set of all possible mappings. There are \\(2N!\\) possible permutations. If it takes at least \\(m\\) times mapping for a certain number to map back to itself, then the number of situations we have is as follows.\n$$ C_{2N}^m \\cdot (m-1)! \\cdot (2N-m)! = \\frac{(2N)!}{m} $$\nThe probability that a number can return to itself after at least m mapping is\n$$ P(m) = \\frac{\\frac{(2N)!}{m}}{(2N)!} = \\frac{1}{m} $$\nWhen \\(m \\ge N+1\\) , various situations are mutually exclusive, so the probability of prisoners\u0026rsquo; success is as follows.\n$$ P = 1 - \\sum_{m=N+1}^{2N} \\frac{1}{m} $$\nApplication and Experiments # Verify the 100 prisoners problem # According to the theoretical analysis in Open the drawer in a way, we can further deduce. If there is a cycle with length greater than \\(50\\), then the number of such cycles must be \\(1\\), so we consider the probability \\(P\u0026rsquo;\\) of the existence of a cycle with length greater than \\(50\\). Obviously, the two events are complementary, so \\(P = 1 - P\u0026rsquo;\\).\nFirst, consider a cycle with length \\(100\\). If the cycle is broken at any point, there are \\(A_{100}^{100}=100!\\) ways. Then, if the cycle is reconnected, any point on the cycle can be considered as the starting point, that is, rotating the cycle is considered the same case, so the actual number of cases is \\(\\frac{100!}{100}\\) . The total number of cases for randomly placing \\(100\\) notes in \\(100\\) drawers is \\(A_{100}^{100}={100!}\\) . Therefore, the probability of there being a cycle with length \\(100\\) is:\n$$ P_{100}=\\frac{\\frac{100!}{100}}{100} = \\frac{1}{100} $$\nSimilarly, consider a cycle with length \\(k (k \u0026gt; 50)\\) . First, select \\(k\\) elements as the elements of the cycle, a total of \\(C_{100}^k\\) ways; shuffle them on the cycle, a total of \\(\\frac{k!}{k}\\) ; the remaining \\(100-k\\) elements are randomly assigned, a total of \\((100-k)!\\) cases; therefore, the number of cases with a cycle of length \\(k\\) is:\n$$ C_{100}^k \\cdot \\frac{k!}{k} \\cdot (100-k)! = \\frac{100!}{k} $$\nTherefore, the probability of there being a cycle of length \\(k\\) is:\n$$ P_k=\\frac{\\frac{100!}{k}}{100!}=\\frac{1}{k} $$\nSince the cases of the existence of a cycle with length \\(k (k \u0026gt; 50)\\) are independent of each other, \\(P\u0026rsquo;=\\sum_{i=51}^{100} P_i\\) , and using a computer program, \\(P\u0026rsquo; \\approx 0.688\\) can be obtained.\nThe C language code is as follows.\n#include \u0026lt;stdio.h\u0026gt; int main() { double sum = 0; for (int i = 51; i \u0026lt;= 100; i++) { sum = sum + 1.0 / i; } printf(\u0026#34;sum=%lf\\n\u0026#34;, sum); return 0; } The results are as follows:\ngcc sum_p.c - o sum_p \u0026amp;\u0026amp;./ sum_p sum = 0.0688172 Therefore, \\(P=1-P\u0026rsquo;\\approx 0.312\\), this method is obviously more feasible than the probability of \\(7.89 \\times 10^{-31}\\) in randomly open the drawer.\nVerify general problem # Verify the calculation results using MATLAB simulation. When the number of prisoners is \\(2-50\\), for each case, simulate \\(10000\\) times and count the frequency of prisoner release to obtain the curve in Figure.\nThe red curve is drawn according to \\(P=1-\\sum_{m=N+1}^{2N} \\frac{1}{m}\\) from General problem, and the blue curve is obtained by simulation. The results obtained from the simulation are roughly the same as the theoretical calculation value, which verifies the correctness of the theoretical calculation.\nThe MATLAB code is as follows.\nloop_times = 10000; theo_prob = []; simu_prob = []; for i = 2:2:50 total = 0; for k = i / 2 + 1:1:i total = total + 1 / k; end theo_prob = [theo_prob, 1 - total]; not_release = 0; for j = 1:1:loop_times draw = rand([i, 1]); [temp, draw] = sort(draw); flag = 0; for n = 1:1:i index = n; for m = 1:1:i index = draw(index); if (index == n) \u0026amp;\u0026amp; (m \u0026lt;= i / 2) break elseif (index == n) \u0026amp;\u0026amp; (m \u0026gt;= i / 2 + 1) flag = 1; n = i + 1; break end end end not_release = not_release + flag; end simu_prob = [simu_prob, 1 - not_release / loop_times]; i; end x = 2:2:50; plot(x, theo_prob, \u0026#39;ro-\u0026#39;) hold on; plot(x, simu_prob, \u0026#39;b+-\u0026#39;) legend({\u0026#39;Theoretical value\u0026#39;, \u0026#39;Analog value\u0026#39;}, \u0026#39;Location\u0026#39;, \u0026#39;southwest\u0026#39;) title(\u0026#34;Probability of prisoners\u0026#39; success\u0026#34;) xlabel(\u0026#34;Number of prisoners\u0026#34;) ylabel(\u0026#34;Probability value\u0026#34;) Scaling the formula \\(P=1-\\sum_{m=N+1}^{2N}\\frac{1}{m}\\) gives the following result.\n$$ P \u0026gt; 1-\\int_{N+1}^{2N} \\frac{1}{dx} = 1-\\ln 2 \\approx 0.307 $$\nIt follows that this probability has a lower bound of \\(1-\\ln2\\).\nConclusion # According to the above calculation results, we know that even if there are more prisoners, all prisoners still have a probability of nearly one third of success - this is contrary to our intuition, at the beginning, our intuition told us that it is almost impossible to let \\(100\\) prisoners find their own numbers. The key point of this strategy is that it combines the results of all people together, trying to succeed or fail together as much as possible. In the end, there are only two possibilities: there is a cycle with length greater than \\(50\\) or there is not. Therefore, there are only two results: complete victory or complete defeat - there is no result where only a few people fail.\nRely on rigorous calculations, not intuition, to solve problems accurately. ","date":"8 January 2023","externalUrl":null,"permalink":"/en/posts/100-prisoners-problem/","section":"Blogs","summary":" The 100 prisoners problem is a counterintuitive problem. It describes a seemingly impossible event: 100 prisoners have a chance to do the same thing, and only when all the prisoners do it can they survive. In fact, there is a reasonable implementation method for this problem to increase its probability by nearly 30 orders of magnitude. Introduction # The 100 prisoners problem is a mathematical problem in probability theory and combinatorics. In this problem, 100 numbered prisoners must find their own numbers in one of 100 drawers in order to survive. The rules state that each prisoner may open only 50 drawers and cannot communicate with other prisoners. Danish computer scientist Peter Bro Miltersen first proposed the problem in 2003. As an upgraded version of this problem, there will be \\(2N\\) prisoners. Their corresponding \\(2N\\) number cards are shuffled and placed in \\(2N\\) drawers. Each prisoner must open at most half of all the drawers and find the corresponding number card for his own number. All prisoners will enter the room separately. The failure of any one prisoner will result in the failure of the entire challenge. What is the maximum probability that the prisoners winning?\n","title":"100 Prisoners Problem","type":"posts"},{"content":"","date":"8 January 2023","externalUrl":null,"permalink":"/en/tags/circular-permutation/","section":"Tags","summary":"","title":"Circular Permutation","type":"tags"},{"content":"","date":"8 January 2023","externalUrl":null,"permalink":"/en/tags/probability/","section":"Tags","summary":"","title":"Probability","type":"tags"},{"content":"","date":"2023-01-08","externalUrl":null,"permalink":"/tags/%E6%A6%82%E7%8E%87/","section":"Tags","summary":"","title":"概率","type":"tags"},{"content":"","date":"2023-01-08","externalUrl":null,"permalink":"/tags/%E7%8E%AF%E6%8E%92%E5%88%97/","section":"Tags","summary":"","title":"环排列","type":"tags"},{"content":"Data source: Breast Cancer Prediction Dataset\nIntroduction # Worldwide, breast cancer is the most common type of cancer in women and the second highest in terms of mortality rates.Diagnosis of breast cancer is performed when an abnormal lump is found (from self-examination or x-ray) or a tiny speck of calcium is seen (on an x-ray). After a suspicious lump is found, the doctor will conduct a diagnosis to determine whether it is cancerous and, if so, whether it has spread to other parts of the body. This breast cancer dataset was obtained from the University of Wisconsin Hospitals, Madison from Dr. William H. Wolberg.\nImport # The first step is import the raw data to the program code.\ndata \u0026lt;- read.csv(\u0026#34;Breast_cancer_data.csv\u0026#34;) Thanks to the selected raw data containing the header and the built-in function(read.csv) of importing CSV files in the R language, we only need a simple line of code to efficiently import files into our program for subsequent processing.\nUsing the head function, we briefly view part of the data.\nhead(data) # mean_radius mean_texture mean_perimeter mean_area mean_smoothness diagnosis # 1 17.99 10.38 122.80 1001.0 0.11840 0 # 2 20.57 17.77 132.90 1326.0 0.08474 0 # 3 19.69 21.25 130.00 1203.0 0.10960 0 # 4 11.42 20.38 77.58 386.1 0.14250 0 # 5 20.29 14.34 135.10 1297.0 0.10030 0 # 6 12.45 15.70 82.57 477.1 0.12780 0 Obviously, each piece of data in the dataset contains 5 data related to breast cancer, and 0 and 1 are used in the last column to indicate whether the diagnosis is confirmed.\nTidy # Generally speaking, in order to analyze the data, we need to clean the data to a certain extent, separate and merge some cell data, discard some unwanted data, and amplify important data, so as to make the data more compact for us to modeling and analyze.\nWe expect to use data to build models. For this data set, every piece of information has its utilization value, so in order to ensure that the amount of data is sufficient, we do not need to process it too much.\nCheck for missing data # We need to judge whether there is any data missing data in the data set to avoid unforeseen errors when building the model.\nsum(is.na(data)) Fortunately, there is no data missing in our dataset, and each piece of data is complete and full of utilization value.\nView redundant data # First of all, we need to ensure that the data in the dataset is not duplicated to prevent interference with subsequent modeling.\nduplicated_count \u0026lt;- sum(duplicated(data)) Output 0, no duplicate data.\nAnalyze data # Our goal is to establish a reasonable prediction model to achieve accurate prediction of breast cancer diagnosis through certain quantifiable data.\nAnalyze the distribution of variables # In order to have a basic understanding of the distribution of parameters, we will map the distribution of each variable separately.\nIn this case, the combination of histogram and density distribution map is the most intuitive and effective.\nSince the longitudinal axis of geom_density() is density estimation, in order to be able to draw the histogram and the density estimation in the same coordinate system, it is necessary to change the longitudinal axis of the histogram to density estimation.\nrel_area \u0026lt;- ggplot(data, aes(x = mean_area, y = ..density..)) + geom_histogram(fill = \u0026#34;blue\u0026#34;, color = \u0026#34;black\u0026#34;, size = 0.2, alpha = 0.2, bins = 30) + geom_density() rel_radius \u0026lt;- ggplot(data, aes(x = mean_radius, y = ..density..)) + geom_histogram(fill = \u0026#34;blue\u0026#34;, color = \u0026#34;black\u0026#34;, size = 0.2, alpha = 0.2, bins = 30) + geom_density() rel_texture \u0026lt;- ggplot(data, aes(x = mean_texture, y = ..density..)) + geom_histogram(fill = \u0026#34;blue\u0026#34;, color = \u0026#34;black\u0026#34;, size = 0.2, alpha = 0.2, bins = 30) + geom_density() rel_smooth \u0026lt;- ggplot(data, aes(x = mean_smoothness, y = ..density..)) + geom_histogram(fill = \u0026#34;blue\u0026#34;, color = \u0026#34;black\u0026#34;, size = 0.2, alpha = 0.2, bins = 30) + geom_density() rel_perimeter \u0026lt;- ggplot(data, aes(x = mean_perimeter, y = ..density..)) + geom_histogram(fill = \u0026#34;blue\u0026#34;, color = \u0026#34;black\u0026#34;, size = 0.2, alpha = 0.2, bins = 30) + geom_density() Finally, we use the grid.arrange function in the gridExtra library to arrange the charts together.\ngrid.arrange(rel_area, rel_radius, rel_texture, rel_smooth, rel_perimeter, nrow = 3, ncol = 2 ) We note that the distribution of all data has two properties:\nAll data is continuously distributed in a certain interval. Each data is distributed in large quantities near a value, and the farther away it is, the less distributed it will be. Analysis of univariate diagnostic results # We analyze the correlation between each variable and the diagnostic results. Here we use the method of drawing a box pattern.\nmean_radius # ggplot(data, aes(x = factor(diagnosis), y = mean_radius)) + geom_boxplot(outlier.colour = \u0026#34;blue\u0026#34;, outlier.shape = 5, outlier.size = 4) + labs(title = \u0026#34;Plot of mean_radius\u0026#34;, x = \u0026#34;diagnosis\u0026#34;) It can be seen that the average tumor radius of patients with breast cancer is mainly between 10 and 15, while the undiagnosed is mainly between 15 and 20.\nmean_texture # ggplot(data, aes(x = factor(diagnosis), y = mean_texture)) + geom_boxplot(outlier.colour = \u0026#34;blue\u0026#34;, outlier.shape = 5, outlier.size = 4) + labs(title = \u0026#34;Plot of mean_texture\u0026#34;, x = \u0026#34;diagnosis\u0026#34;) It can be found that the average texture value of the patient\u0026rsquo;s tumor is between 15 and 20, while that has not been diagnosed is between 20 and 25.\nmean_perimeter # ggplot(data, aes(x = factor(diagnosis), y = mean_perimeter)) + geom_boxplot(outlier.colour = \u0026#34;blue\u0026#34;, outlier.shape = 5, outlier.size = 4) + labs(title = \u0026#34;Plot of mean_perimeter\u0026#34;, x = \u0026#34;diagnosis\u0026#34;) As can be seen from the figure, the average perimeter of tumors in confirmed patients is between 70 and 90, while those that have not been diagnosed is between 100 and 130.\nmean_area # ggplot(data, aes(x = factor(diagnosis), y = mean_area)) + geom_boxplot(outlier.colour = \u0026#34;blue\u0026#34;, outlier.shape = 5, outlier.size = 4) + labs(title = \u0026#34;Plot of mean_area\u0026#34;, x = \u0026#34;diagnosis\u0026#34;) It can be found that the average tumor area of confirmed patients is about 500, while the main undiagnosed tumors are between 750 and 1250.\nmean_smoothness # ggplot(data, aes(x = factor(diagnosis), y = mean_smoothness)) + geom_boxplot(outlier.colour = \u0026#34;blue\u0026#34;, outlier.shape = 5, outlier.size = 4) + labs(title = \u0026#34;Plot of mean_smoothness\u0026#34;, x = \u0026#34;diagnosis\u0026#34;) We have noticed that there is an intersection of tumor smoothness in patients with positive or negative confirmed results, but in general, the value of confirmed patients will be lower.\nRelevance analysis # According to the general process of modeling, we need to carry out correlation analysis of each variable.\nIf the correlation between variables is significant, it will affect the predictive effect of the model.\ncor_analysis \u0026lt;- cor(data[c(1:5)]) corrplot(cor_analysis, method = \u0026#34;number\u0026#34;) Through correlation analysis, we found that the relationship between the three variables is very significant.\nThey are radius, perimeter and area.\nThese three values are obviously highly correlated, so we need to filter them.\nCorrelation between variables and diagnostic results # Before the final accuracy of the qualitative analysis model, we want to use the image to see the relationship between variables and diagnostic results first. There are some intuitive impressions.\nWe have a total of five variables:\nmean_radius mean_texture mean_perimeter mean_area mean_smoothness The correlation between them and diagnostic results is combined, and there are a total of 10 situations that need to be discussed and analyzed.\nThe red dot in the chart below indicates that the diagnostic result is undiagnosed.\nWeak correlation of variables # radius \u0026amp; texture, radius \u0026amp; smoothness, texture \u0026amp; perimeter, texture \u0026amp; area, texture \u0026amp; smoothness, perimeter \u0026amp; smoothness, area \u0026amp; smoothness\nWe can have an intuitive understanding of the relationship between the two variables through the scatter distribution in the chart.\nStrong correlation of variables # radius \u0026amp; perimeter, radius \u0026amp; area, perimeter \u0026amp; area\n# radius \u0026amp; perimeter rpPlot \u0026lt;- ggplot(data, aes( x = mean_radius, y = mean_perimeter, color = as.factor(diagnosis) )) + geom_point() + theme_minimal() + theme(legend.position = \u0026#34;none\u0026#34;) # radius \u0026amp; area raPlot \u0026lt;- ggplot(data, aes( x = mean_radius, y = mean_area, color = as.factor(diagnosis) )) + geom_point() + theme_minimal() + theme(legend.position = \u0026#34;none\u0026#34;) # perimeter \u0026amp; area paPlot \u0026lt;- ggplot(data, aes( x = mean_perimeter, y = mean_area, color = as.factor(diagnosis) )) + geom_point() + theme_minimal() + theme(legend.position = \u0026#34;none\u0026#34;) grid.arrange(rpPlot, raPlot, paPlot, nrow = 1, ncol = 3 ) As previously expected, they are quite relevant, and their analysis can be carried out at the same time.\nObviously, from the chart, patients with an average tumor radius of 10 to 15 have a higher probability of being diagnosed with breast cancer.\nModel # Dataset segmentation # For subsequent modeling, we need to randomly divide the data set into two parts, one to train the prediction model, and the other to test the accuracy of the model.\nset.seed(123) # Set the repeatability set.seed() to ensure that it is repeatable train \u0026lt;- sample(nrow(data), 0.7 * nrow(data)) train_data \u0026lt;- data[train, ] test_data \u0026lt;- data[-train, ] train_ data means training data, validate_ data stands for inspection data\nLogistic regression modeling # Since the final prediction results are 0 and 1, it is not suitable to use linear regression.\nHere we choose to use the idea of logical regression to build a model.\nThe connection function used in logical regression is the best representative of the Sigmoid function, that is, the logistic function.\nFrom the above analysis, we select radius among radius, perimeter and area for modeling.\nmodel \u0026lt;- glm( data = train_data, formula = diagnosis ~ mean_texture + mean_smoothness + mean_radius, family = binomial(link = \u0026#34;logit\u0026#34;) ) model \u0026lt;- step(model) # Carry out the step-by-step regression method for data analysis summary(model) Use summary (model) to view the model.\nCall: glm(formula = diagnosis ~ mean_texture + mean_smoothness + mean_radius, family = binomial(link = \u0026#34;logit\u0026#34;), data = train_data) Deviance Residuals: Min 1Q Median 3Q Max -2.91948 -0.03436 0.04781 0.21133 2.01672 Coefficients: Estimate Std. Error z value Pr(\u0026gt;|z|) (Intercept) 40.52957 5.09828 7.950 1.87e-15 *** mean_texture -0.34187 0.06622 -5.163 2.44e-07 *** mean_smoothness -140.35265 21.54005 -6.516 7.23e-11 *** mean_radius -1.36821 0.17827 -7.675 1.65e-14 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 (Dispersion parameter for binomial family taken to be 1) Null deviance: 515.0 on 397 degrees of freedom Residual deviance: 134.9 on 394 degrees of freedom AIC: 142.9 Number of Fisher Scoring iterations: 8 From the results of summary, it can be seen that the three variables we selected contribute significantly to the results.\nCalculate the fitting threshold # Here we use roc from pRoc to find the optimal threshold.\npre \u0026lt;- predict(model, type = \u0026#34;response\u0026#34;, train_data) modelroc \u0026lt;- roc(train_data$diagnosis, pre) plot(modelroc, print.auc = TRUE, auc.polygon = TRUE, grid = c(0.1, 0.2), grid.col = c(\u0026#34;green\u0026#34;, \u0026#34;red\u0026#34;), max.auc.polygon = TRUE, auc.polygon.col = \u0026#34;skyblue\u0026#34;, print.thres = TRUE ) It can be seen that 0.588 is the threshold we need. Considering it, we choose 0.6 as the threshold of the model.\nTest Set Data Validation # After successfully using the training set data to build a model, we should also use the check data set to check the accuracy of model prediction.\nSince the forecast is a number, and what we ultimately want is a confirmed or undiagnosed result, we need a threshold to classify the predicted value to get a numerical result of 1 or 0.\ntest_data$prob \u0026lt;- model %\u0026gt;% predict(type = \u0026#34;response\u0026#34;, newdata = test_data) test_data$prob \u0026lt;- ifelse(test_data$prob \u0026gt; 0.6, 1, 0) test_data$diff \u0026lt;- ifelse(test_data$diagnosis == test_data$prob, 1, 0) Evaluate the predictive effect of the model # We can draw a pie chart to visualize the accuracy of the model.\ndiff_count \u0026lt;- test_data %\u0026gt;% count(diff, name = \u0026#34;count\u0026#34;) diff_count$diff \u0026lt;- ifelse(diff_count$diff == 0, \u0026#34;False\u0026#34;, \u0026#34;True\u0026#34;) diff_plot \u0026lt;- diff_count %\u0026gt;% ggplot(mapping = aes( x = 1, y = count, fill = factor(diff), )) diff_plot + geom_bar(stat = \u0026#34;identity\u0026#34;) + coord_polar(theta = \u0026#34;y\u0026#34;) + scale_x_continuous(name = NULL, breaks = NULL) + scale_y_continuous(name = NULL, breaks = NULL) + labs( x = \u0026#34;\u0026#34;, y = \u0026#34;\u0026#34;, fill = \u0026#34;\u0026#34;, title = \u0026#34;Model prediction accuracy\u0026#34;, ) + theme( legend.position = \u0026#34;top\u0026#34;, plot.title = element_text(hjust = 0.5, size = 14), # title position ) paste( round(100 * diff_count$count[2] / (diff_count$count[1] + diff_count$count[2]), 2), \u0026#34;%\u0026#34; ) [1] \u0026#34;93.57 %\u0026#34; The result was very surprising. The prediction accuracy of our model was extremely high. In order to avoid errors and remove preset random values, I tried many times and got better results of more than 90%, which shows that our model is stable and accurate.\nFull code library(tidyverse) library(ggplot2) library(corrplot) library(pROC) library(gridExtra) # Import the original data data \u0026lt;- read.csv(\u0026#34;Breast_cancer_data.csv\u0026#34;) head(data) # Check for missing data sum(is.na(data)) # Redundant data view duplicated_count \u0026lt;- sum(duplicated(data)) # Analyze the distribution of each variable rel_area \u0026lt;- ggplot(data, aes(x = mean_area, y = ..density..)) + geom_histogram(fill = \u0026#34;blue\u0026#34;, color = \u0026#34;black\u0026#34;, size = 0.2, alpha = 0.2, bins = 30) + # nolint geom_density() rel_radius \u0026lt;- ggplot(data, aes(x = mean_radius, y = ..density..)) + geom_histogram(fill = \u0026#34;blue\u0026#34;, color = \u0026#34;black\u0026#34;, size = 0.2, alpha = 0.2, bins = 30) + # nolint geom_density() rel_texture \u0026lt;- ggplot(data, aes(x = mean_texture, y = ..density..)) + geom_histogram(fill = \u0026#34;blue\u0026#34;, color = \u0026#34;black\u0026#34;, size = 0.2, alpha = 0.2, bins = 30) + # nolint geom_density() rel_smooth \u0026lt;- ggplot(data, aes(x = mean_smoothness, y = ..density..)) + geom_histogram(fill = \u0026#34;blue\u0026#34;, color = \u0026#34;black\u0026#34;, size = 0.2, alpha = 0.2, bins = 30) + # nolint geom_density() rel_perimeter \u0026lt;- ggplot(data, aes(x = mean_perimeter, y = ..density..)) + geom_histogram(fill = \u0026#34;blue\u0026#34;, color = \u0026#34;black\u0026#34;, size = 0.2, alpha = 0.2, bins = 30) + # nolint geom_density() grid.arrange(rel_area, rel_radius, rel_texture, rel_smooth, rel_perimeter, nrow = 3, ncol = 2 ) # Univariate box pattern analysis # mean_radius ggplot(data, aes(x = factor(diagnosis), y = mean_radius)) + geom_boxplot(outlier.colour = \u0026#34;blue\u0026#34;, outlier.shape = 5, outlier.size = 4) + labs(title = \u0026#34;Plot of mean_radius\u0026#34;, x = \u0026#34;diagnosis\u0026#34;) # mean_texture ggplot(data, aes(x = factor(diagnosis), y = mean_texture)) + geom_boxplot(outlier.colour = \u0026#34;blue\u0026#34;, outlier.shape = 5, outlier.size = 4) + labs(title = \u0026#34;Plot of mean_texture\u0026#34;, x = \u0026#34;diagnosis\u0026#34;) # mean_perimeter ggplot(data, aes(x = factor(diagnosis), y = mean_perimeter)) + geom_boxplot(outlier.colour = \u0026#34;blue\u0026#34;, outlier.shape = 5, outlier.size = 4) + labs(title = \u0026#34;Plot of mean_perimeter\u0026#34;, x = \u0026#34;diagnosis\u0026#34;) # mean_area ggplot(data, aes(x = factor(diagnosis), y = mean_area)) + geom_boxplot(outlier.colour = \u0026#34;blue\u0026#34;, outlier.shape = 5, outlier.size = 4) + labs(title = \u0026#34;Plot of mean_area\u0026#34;, x = \u0026#34;diagnosis\u0026#34;) # mean_smoothness ggplot(data, aes(x = factor(diagnosis), y = mean_smoothness)) + geom_boxplot(outlier.colour = \u0026#34;blue\u0026#34;, outlier.shape = 5, outlier.size = 4) + labs(title = \u0026#34;Plot of mean_smoothness\u0026#34;, x = \u0026#34;diagnosis\u0026#34;) # Correlation analysis between variables cor_analysis \u0026lt;- cor(data[c(1:5)]) corrplot(cor_analysis, method = \u0026#34;number\u0026#34;) # Relevance between variables and results # ----------------------------Related variables------------------------------ # # radius \u0026amp; texture rtPlot \u0026lt;- ggplot(data, aes( x = mean_radius, y = mean_texture, color = as.factor(diagnosis) )) + geom_point() + theme_minimal() + theme(legend.position = \u0026#34;none\u0026#34;) # radius \u0026amp; smoothness rsPlot \u0026lt;- ggplot(data, aes( x = mean_radius, y = mean_smoothness, color = as.factor(diagnosis) )) + geom_point() + theme_minimal() + theme(legend.position = \u0026#34;none\u0026#34;) # texture \u0026amp; perimeter tpPlot \u0026lt;- ggplot(data, aes( x = mean_texture, y = mean_perimeter, color = as.factor(diagnosis) )) + geom_point() + theme_minimal() + theme(legend.position = \u0026#34;none\u0026#34;) # texture \u0026amp; area taPlot \u0026lt;- ggplot(data, aes( x = mean_texture, y = mean_area, color = as.factor(diagnosis) )) + geom_point() + theme_minimal() + theme(legend.position = \u0026#34;none\u0026#34;) # texture \u0026amp; smoothness tsPlot \u0026lt;- ggplot(data, aes( x = mean_texture, y = mean_smoothness, color = as.factor(diagnosis) )) + geom_point() + theme_minimal() + theme(legend.position = \u0026#34;none\u0026#34;) # perimeter \u0026amp; smoothness psPlot \u0026lt;- ggplot(data, aes( x = mean_perimeter, y = mean_smoothness, color = as.factor(diagnosis) )) + geom_point() + theme_minimal() + theme(legend.position = \u0026#34;none\u0026#34;) # area \u0026amp; smoothness asPlot \u0026lt;- ggplot(data, aes( x = mean_area, y = mean_smoothness, color = as.factor(diagnosis) )) + geom_point() + theme_minimal() + theme(legend.position = \u0026#34;none\u0026#34;) grid.arrange(rtPlot, rsPlot, tpPlot, taPlot, tsPlot, psPlot, asPlot, nrow = 3, ncol = 3 ) # ---------------------------Strong related variable---------------------------- # # radius \u0026amp; perimeter rpPlot \u0026lt;- ggplot(data, aes( x = mean_radius, y = mean_perimeter, color = as.factor(diagnosis) )) + geom_point() + theme_minimal() + theme(legend.position = \u0026#34;none\u0026#34;) # radius \u0026amp; area raPlot \u0026lt;- ggplot(data, aes( x = mean_radius, y = mean_area, color = as.factor(diagnosis) )) + geom_point() + theme_minimal() + theme(legend.position = \u0026#34;none\u0026#34;) # perimeter \u0026amp; area paPlot \u0026lt;- ggplot(data, aes( x = mean_perimeter, y = mean_area, color = as.factor(diagnosis) )) + geom_point() + theme_minimal() + theme(legend.position = \u0026#34;none\u0026#34;) grid.arrange(rpPlot, raPlot, paPlot, nrow = 1, ncol = 3 ) # Dataset segmentation set.seed(123) # Set the repeatable set.seed() to ensure that it is repeatable train \u0026lt;- sample(nrow(data), 0.7 * nrow(data)) train_data \u0026lt;- data[train, ] test_data \u0026lt;- data[-train, ] # Logical regression modeling model \u0026lt;- glm( data = train_data, formula = diagnosis ~ mean_texture + mean_smoothness + mean_radius, family = binomial(link = \u0026#34;logit\u0026#34;) ) model \u0026lt;- step(model) # step-by-step regression method summary(model) # Export all results # Calculate the fitting threshold pre \u0026lt;- predict(model, type = \u0026#34;response\u0026#34;, train_data) modelroc \u0026lt;- roc(train_data$diagnosis, pre) plot(modelroc, print.auc = TRUE, auc.polygon = TRUE, grid = c(0.1, 0.2), grid.col = c(\u0026#34;green\u0026#34;, \u0026#34;red\u0026#34;), max.auc.polygon = TRUE, auc.polygon.col = \u0026#34;skyblue\u0026#34;, print.thres = TRUE ) # Test set data verification test_data$prob \u0026lt;- model %\u0026gt;% predict(type = \u0026#34;response\u0026#34;, newdata = test_data) test_data$prob \u0026lt;- ifelse(test_data$prob \u0026gt; 0.6, 1, 0) test_data$diff \u0026lt;- ifelse(test_data$diagnosis == test_data$prob, 1, 0) # Test set data forecast statistics percentage diff_count \u0026lt;- test_data %\u0026gt;% count(diff, name = \u0026#34;count\u0026#34;) diff_count$diff \u0026lt;- ifelse(diff_count$diff == 0, \u0026#34;False\u0026#34;, \u0026#34;True\u0026#34;) diff_plot \u0026lt;- diff_count %\u0026gt;% ggplot(mapping = aes( x = 1, y = count, fill = factor(diff), )) diff_plot + geom_bar(stat = \u0026#34;identity\u0026#34;) + coord_polar(theta = \u0026#34;y\u0026#34;) + scale_x_continuous(name = NULL, breaks = NULL) + scale_y_continuous(name = NULL, breaks = NULL) + labs( x = \u0026#34;\u0026#34;, y = \u0026#34;\u0026#34;, fill = \u0026#34;\u0026#34;, title = \u0026#34;Model prediction accuracy\u0026#34;, ) + theme( legend.position = \u0026#34;top\u0026#34;, plot.title = element_text(hjust = 0.5, size = 14), ) paste( round(100 * diff_count$count[2] / (diff_count$count[1] + diff_count$count[2]), 2), \u0026#34;%\u0026#34; ) ","date":"22 June 2022","externalUrl":null,"permalink":"/en/posts/breast-cancer-prediction/","section":"Blogs","summary":"Data source: Breast Cancer Prediction Dataset\nIntroduction # Worldwide, breast cancer is the most common type of cancer in women and the second highest in terms of mortality rates.Diagnosis of breast cancer is performed when an abnormal lump is found (from self-examination or x-ray) or a tiny speck of calcium is seen (on an x-ray). After a suspicious lump is found, the doctor will conduct a diagnosis to determine whether it is cancerous and, if so, whether it has spread to other parts of the body. ","title":"Breast Cancer Prediction","type":"posts"},{"content":"","date":"22 June 2022","externalUrl":null,"permalink":"/en/tags/data-analysis/","section":"Tags","summary":"","title":"Data Analysis","type":"tags"},{"content":"","date":"22 June 2022","externalUrl":null,"permalink":"/en/series/data-analysis-in-r/","section":"Series","summary":"","title":"Data Analysis in R","type":"series"},{"content":"","date":"22 June 2022","externalUrl":null,"permalink":"/en/tags/r-language/","section":"Tags","summary":"","title":"R Language","type":"tags"},{"content":"","date":"2022-06-22","externalUrl":null,"permalink":"/tags/r%E8%AF%AD%E8%A8%80/","section":"Tags","summary":"","title":"R语言","type":"tags"},{"content":"","date":"2022-06-22","externalUrl":null,"permalink":"/series/r%E8%AF%AD%E8%A8%80%E6%95%B0%E6%8D%AE%E5%88%86%E6%9E%90/","section":"Series","summary":"","title":"R语言数据分析","type":"series"},{"content":"","date":"2022-06-22","externalUrl":null,"permalink":"/tags/%E6%95%B0%E6%8D%AE%E5%88%86%E6%9E%90/","section":"Tags","summary":"","title":"数据分析","type":"tags"},{"content":"Data sources:Jail deaths in America: data and key findings of Dying Inside\nIt contains data on large prisons with 750 or more prisoners in the United States. In order to ensure that it checked the number of deaths across the country, it also included data on the 10 largest prisons in each state.\nThese data cover 523 prisons or prison systems. The data covers the period from 2008 to 2019.\nIntroduction # The U.S. government does not release jail by jail mortality data, keeping the public and policy makers in the dark about facilities with high rates of death. In a first-of-its-kind accounting, Reuters obtained and is releasing that data to the public.\nWhat if the jail in your community had an outsized death rate, but no one knew? For decades, communities across the country have faced that quandary. The Justice Department collects jail death data, but locks the information away, leaving policymakers, investigators and activists unaware of problem facilities.\nReuters journalists filed more than 1,500 public records requests to gain death data from 2008 to 2019 in the nation’s biggest jails. Today, jail by jail and state by state, it is making that information available to the public. Reuters examined every large jail in the United States, those with 750 or more inmates. And, to ensure it examined deaths across the country, it obtained data for the 10 largest jails in each state. The data covers 523 jails or jail systems.\nDeaths in Large US Prisons # Used librarys: tidyverse, ggplot2\nImport # The first step is import the raw data to the program code.\ndata_csv \u0026lt;- read.csv(\u0026#34;all_deaths.csv\u0026#34;) Selection # We want to have an intuitive understanding of the data as a whole. At this time, we don\u0026rsquo;t need specific details, so we only filter out the year data.\ndeath_count \u0026lt;- data_csv %\u0026gt;% select(year) %\u0026gt;% count(year, name = \u0026#34;count\u0026#34;) %\u0026gt;% mutate(year = factor(year, levels = seq(2008, 20019, by = 1))) Note that in this code snippet, we convert year to a factor so that the x-axis coordinates of the histogram are scattered year data, using the mutata function.\nPlot # We chose ggplot2 to draw exquisite and intuitive data analysis images.\nWe use year as the parameter of the x-axis. Since we have previously calculated the occurrence frequency of each year, we only need to specify X and Y.\ndataplot \u0026lt;- death_count %\u0026gt;% ggplot(mapping = aes( x = year, y = count, )) geom_col() Since the data display is not intuitive, we add color to it and according to the principle that the larger the data, the darker the color. Because we don\u0026rsquo;t need to annotate the color, we add the guides(fill = FALSE) statement. Then we add a specific number display to each bar diagram to make the image intuitive and detailed. dataplot \u0026lt;- death_count %\u0026gt;% ggplot(mapping = aes( x = year, y = count, fill = -log(count), )) + guides(fill = FALSE) + geom_col() + geom_text(mapping = aes( label = count, )) Add auxiliary information # Finally, we add auxiliary information to expand the chart, including the title, subtitle, and data source.\nThen fine-tune the title position and select the appropriate font family to further upgrade the image effect.\ndataplot + labs( x = NULL, y = \u0026#34;Death toll\u0026#34;, title = \u0026#34;Total number of deaths in large prisons in the United States\u0026#34;, subtitle = \u0026#34;death data from 2008 to 2019 in the nation\u0026#39;s biggest jails\u0026#34;, caption = \u0026#34;Data sources: https://www.reuters.com/investigates/special-report/usa-jails-graphic/\u0026#34;, # nolint ) + theme( plot.title = element_text(hjust = 0.5, size = 14), # title position text = element_text(family = \u0026#34;JetBrains Mono\u0026#34;), # font ) Conclusions of the study # Between 2008 and 2019, the number of deaths in large prisons in the United States gradually increased, from 469 to a maximum of 731, and gradually stabilized at about 700.\nFull code library(tidyverse) library(ggplot2) # Import data_csv \u0026lt;- read.csv(\u0026#34;all_deaths.csv\u0026#34;) # data select death_count \u0026lt;- data_csv %\u0026gt;% select(year) %\u0026gt;% count(year, name = \u0026#34;count\u0026#34;) %\u0026gt;% mutate(year = factor(year, levels = seq(2008, 20019, by = 1))) # Plot dataplot \u0026lt;- death_count %\u0026gt;% ggplot(mapping = aes( x = year, y = count, fill = -log(count), )) + # guides(fill = FALSE) + geom_col() + geom_text(mapping = aes( label = count, )) # Add auxiliary information. dataplot + labs( x = NULL, y = \u0026#34;Death toll\u0026#34;, title = \u0026#34;Total number of deaths in large prisons in the United States\u0026#34;, subtitle = \u0026#34;death data from 2008 to 2019 in the nation\u0026#39;s biggest jails\u0026#34;, caption = \u0026#34;Data sources: https://www.reuters.com/investigates/special-report/usa-jails-graphic/\u0026#34;, # nolint ) + theme( legend.position = \u0026#34;none\u0026#34;, plot.title = element_text(hjust = 0.5, size = 14), # title position text = element_text(family = \u0026#34;JetBrains Mono\u0026#34;), # font ) Prison deaths by state # In this part, we try to further explore the details of prison deaths in the United States on a state by state basis.\nImport # The first step is import the raw data to the program code.\ndata_csv \u0026lt;- read.csv(\u0026#34;all_deaths.csv\u0026#34;) Process raw data # Since we need the total number of deaths in all prisons in each state each year, we first use table to calculate the same information.\nThen we will change the table format to a data box form again, and change the index name to facilitate our subsequent plotting.\ndata_csv \u0026lt;- table(data_csv$year, data_csv$state) %\u0026gt;% as.data.frame() %\u0026gt;% rename(year = Var1, state = Var2, count = Freq) Plot # Based on the processing of the above data, we can efficiently make a broken line chart of the number of deaths in large prisons in various states on an image to facilitate follow-up analysis.\ndata_plot \u0026lt;- data_csv %\u0026gt;% ggplot(mapping = aes( x = year, y = count, group = state, color = state, )) data_plot + geom_line() + labs( x = NULL, y = \u0026#34;Death toll\u0026#34;, title = \u0026#34;Prison deaths by state in the United States\u0026#34;, subtitle = \u0026#34;death data from 2008 to 2019 in the nation\u0026#39;s biggest jails\u0026#34;, caption = \u0026#34;Data sources: https://www.reuters.com/investigates/special-report/usa-jails-graphic/\u0026#34;, # nolint ) + theme( plot.title = element_text(hjust = 0.6, size = 14), # title position text = element_text(family = \u0026#34;JetBrains Mono\u0026#34;), # font ) Conclusions of the study # We can clearly see that most of the data curves are gathered in one place, speculating because the prison is similar in size. The data from four states is very eye-catching. California Florida Texas Pennsylvania The data of the four states mentioned above far exceeds that of other states. It is reasonable to speculate that there may be more and larger prisons, and there may be relatively high crime rates. Full code library(tidyverse) library(ggplot2) data_csv \u0026lt;- read.csv(\u0026#34;all_deaths.csv\u0026#34;) data_csv \u0026lt;- table(data_csv$year, data_csv$state) %\u0026gt;% as.data.frame() %\u0026gt;% rename(year = Var1, state = Var2, count = Freq) data_plot \u0026lt;- data_csv %\u0026gt;% ggplot(mapping = aes( x = year, y = count, group = state, color = state, )) data_plot + geom_line() + labs( x = NULL, y = \u0026#34;Death toll\u0026#34;, title = \u0026#34;Prison deaths by state in the United States\u0026#34;, subtitle = \u0026#34;death data from 2008 to 2019 in the nation\u0026#39;s biggest jails\u0026#34;, caption = \u0026#34;Data sources: https://www.reuters.com/investigates/special-report/usa-jails-graphic/\u0026#34;, # nolint ) + theme( plot.title = element_text(hjust = 0.6, size = 14), # title position text = element_text(family = \u0026#34;JetBrains Mono\u0026#34;), # font ) ","date":"21 June 2022","externalUrl":null,"permalink":"/en/posts/jail-deaths-in-america/","section":"Blogs","summary":"Data sources:Jail deaths in America: data and key findings of Dying Inside\nIt contains data on large prisons with 750 or more prisoners in the United States. In order to ensure that it checked the number of deaths across the country, it also included data on the 10 largest prisons in each state.\n","title":"Jail Deaths In America","type":"posts"},{"content":"Data source: Top 50 Most Followed Twitter Accounts\nIntroduction # The data lists the top 50 most-followed accounts on Twitter, with each total rounded to the nearest hundred thousand, as well as the profession or activity of each user. Account totals and monthly changes in the ranking were last updated on May 12, 2022.\nFan leaderboard # We try to use this data set to make intuitive charts of the top 50 accounts of Twitter fans.\nImport # The first step is import the raw data to the program code.\ndata_csv \u0026lt;- read.csv(\u0026#34;Top 50 Most Followed Twitter Accounts.csv\u0026#34;) Visualize # Due to the non-repeatable nature of the Twitter account id, we naturally chose the user ID as the y-axis data. We reorder the accounts with the number of fans and display them on the image from more to less. In order to make the image comparison more intuitive, we creatively use the number of fans to draw gradient colors, from deep to light represents the number of fans from more to less . dataplot \u0026lt;- data_csv %\u0026gt;% ggplot(mapping = aes( x = Followers..millions., y = reorder(Account.username, Followers..millions.), fill = -log(Followers..millions.), )) + geom_bar( stat = \u0026#34;identity\u0026#34;, ) + guides(fill = \u0026#34;none\u0026#34;) + geom_text(mapping = aes( label = Followers..millions., )) # Add auxiliary information. dataplot + labs( x = \u0026#34;Followers (Millions)\u0026#34;, y = \u0026#34;Username\u0026#34;, title = \u0026#34;Top 50 in Twitter\u0026#34;, subtitle = \u0026#34;Information was last updated on May 12, 2022.\u0026#34;, caption = \u0026#34;Data sources: https://www.kaggle.com/datasets/hassanshehzadk/top-50-most-followed-twitter-accounts?resource=download\u0026#34;, ) + theme( plot.title = element_text(hjust = 0.4, size = 14), # title position panel.grid.minor = element_blank(), # Secondary grid lines text = element_text(family = \u0026#34;Hack Nerd Font\u0026#34;), # font axis.text.x = element_text(angle = 45, vjust = 1, hjust = 1) ) Full code library(tidyverse) library(ggplot2) # Import data_csv \u0026lt;- read.csv(\u0026#34;Top 50 Most Followed Twitter Accounts.csv\u0026#34;) # Plot dataplot \u0026lt;- data_csv %\u0026gt;% ggplot(mapping = aes( x = Followers..millions., y = reorder(Account.username, Followers..millions.), fill = -log(Followers..millions.), )) + geom_bar( stat = \u0026#34;identity\u0026#34;, ) + guides(fill = \u0026#34;none\u0026#34;) + geom_text(mapping = aes( label = Followers..millions., )) # Add auxiliary information. dataplot + labs( x = \u0026#34;Followers (Millions)\u0026#34;, y = \u0026#34;Username\u0026#34;, title = \u0026#34;Top 50 in Twitter\u0026#34;, subtitle = \u0026#34;Information was last updated on May 12, 2022.\u0026#34;, caption = \u0026#34;Data sources: https://www.kaggle.com/datasets/hassanshehzadk/top-50-most-followed-twitter-accounts?resource=download\u0026#34;, # nolint ) + theme( plot.title = element_text(hjust = 0.4, size = 14), # title position panel.grid.minor = element_blank(), # Secondary grid lines text = element_text(family = \u0026#34;Hack Nerd Font\u0026#34;), # font axis.text.x = element_text(angle = 45, vjust = 1, hjust = 1) ) Account attribution analysis # We also want to see which countries the accounts with a large number of fans come from, so we want to draw a pie chart to check their distribution.\nImport # The first step is import the raw data to the program code.\ndata_csv \u0026lt;- read.csv(\u0026#34;Top 50 Most Followed Twitter Accounts.csv\u0026#34;) Selection # We quickly count the frequency of each country in the data set.\narea_count \u0026lt;- data_csv %\u0026gt;% count(Country, name = \u0026#34;count\u0026#34;) Visualize # Since there is no built-in pie chart drawing method in ggplot, we use geom_bar and coord_polar to try to achieve the same effect. dataplot \u0026lt;- area_count %\u0026gt;% ggplot(mapping = aes( x = 1, y = count, fill = Country, )) dataplot + geom_bar(stat = \u0026#34;identity\u0026#34;) + coord_polar(theta = \u0026#34;y\u0026#34;) + scale_x_continuous(name = NULL, breaks = NULL) + scale_y_continuous(name = NULL, breaks = NULL) + scale_fill_viridis_d(option = \u0026#34;inferno\u0026#34;) Conclusions of the study # Obviously, as a native national software of the United States, and considering the large population base of the United States, the number of countries whose accounts belong to the United States far exceeds others. India, as a populous country, unexpectedly became the second place after the United States. The number of other countries is basically the same. Full code library(tidyverse) library(ggplot2) data_csv \u0026lt;- read.csv(\u0026#34;Top 50 Most Followed Twitter Accounts.csv\u0026#34;) area_count \u0026lt;- data_csv %\u0026gt;% count(Country, name = \u0026#34;count\u0026#34;) dataplot \u0026lt;- area_count %\u0026gt;% ggplot(mapping = aes( x = 1, y = count, fill = Country, )) dataplot + geom_bar(stat = \u0026#34;identity\u0026#34;) + coord_polar(theta = \u0026#34;y\u0026#34;) + scale_x_continuous(name = NULL, breaks = NULL) + scale_y_continuous(name = NULL, breaks = NULL) + scale_fill_viridis_d(option = \u0026#34;inferno\u0026#34;) + labs( x = \u0026#34;Followers (Millions)\u0026#34;, y = \u0026#34;Username\u0026#34;, fill = \u0026#34;Country\u0026#34;, title = \u0026#34;Country of Account\u0026#34;, subtitle = \u0026#34;Calculate the top 50 fan accounts on Twitter.\\nInformation was last updated on May 12, 2022.\u0026#34;, # nolint caption = \u0026#34;Data sources: https://www.kaggle.com/datasets/hassanshehzadk/top-50-most-followed-twitter-accounts?resource=download\u0026#34;, # nolint ) + theme( plot.title = element_text(hjust = 0.6, size = 14), # title position axis.text.x = element_text(angle = 45, vjust = 1, hjust = 1), plot.caption = element_text(hjust = 0.3), ) ","date":"20 June 2022","externalUrl":null,"permalink":"/en/posts/top-50-twitter-accounts/","section":"Blogs","summary":"Data source: Top 50 Most Followed Twitter Accounts\nIntroduction # The data lists the top 50 most-followed accounts on Twitter, with each total rounded to the nearest hundred thousand, as well as the profession or activity of each user. Account totals and monthly changes in the ranking were last updated on May 12, 2022.\n","title":"Top 50 Most Followed Twitter Accounts","type":"posts"},{"content":"Data source: Carbon Dioxide Emission Estimates\nIntroduction # The data comes from undata. Last updated on November 1, 2021. It contains the total carbon emissions and per capita carbon emissions of various countries in the world for several years.\nDue to the large number of countries, in this data analysis, we only selected nine countries, namely, China, the United States, France, Germany, Canada, Italy, Japan, Britain and Austria.\nDraw in a picture # We hope to have an intuitive understanding of the data of these nine countries, so we chose to draw it in the same chart.\nImport # The first step is import the raw data to the program code.\ndata_csv \u0026lt;- read.csv(\u0026#34;SYB64_310_202110_Carbon Dioxide Emission Estimates.csv\u0026#34;) Tidy # We need to filter the data of these nine countries from the data set, and in order to be able to compare, we only need per capita data, not overall data.\nworld_data \u0026lt;- data_csv %\u0026gt;% filter(Country %in% c(\u0026#34;China\u0026#34;, \u0026#34;United States of America\u0026#34;, \u0026#34;United Kingdom\u0026#34;, \u0026#34;France\u0026#34;, \u0026#34;Germany\u0026#34;, \u0026#34;Australia\u0026#34;, \u0026#34;Japan\u0026#34;, \u0026#34;Canada\u0026#34;, \u0026#34;Italy\u0026#34;)) %\u0026gt;% filter(Series == \u0026#34;Emissions per capita (metric tons of carbon dioxide)\u0026#34;) Visualize # You only need to map the data set by country, by year as x coordinates, and per capita emissions as y coordinates.\nworld_plot \u0026lt;- world_data %\u0026gt;% ggplot(mapping = aes( x = factor(Year), y = as.numeric(Value), color = Country, fill = Country, group = Country, )) world_plot + geom_point() + geom_line() + labs( x = NULL, y = \u0026#34;Emissions per capita (metric tons of carbon dioxide)\u0026#34;, title = \u0026#34;Per capita emissions of some countries\u0026#34;, subtitle = \u0026#34;Data from 1975 to 2018\u0026#34;, caption = \u0026#34;Data sources: http://data.un.org/default.aspx\u0026#34;, ) + theme( plot.title = element_text(size = 14), # title position text = element_text(family = \u0026#34;JetBrains Mono\u0026#34;), # font ) Conclusions of the study # Generally speaking, the emissions of each country have tended to be flat and stable since 2015. Since 1975, the per capita carbon emissions of various countries have increased and decreased, and have tended to be stable. On the whole, the per capita carbon emissions of Australia, Canada and the United States far exceed those of other countries. Full code library(tidyverse) library(ggplot2) data_csv \u0026lt;- read.csv(\u0026#34;SYB64_310_202110_Carbon Dioxide Emission Estimates.csv\u0026#34;) world_data \u0026lt;- data_csv %\u0026gt;% filter(Country %in% c(\u0026#34;China\u0026#34;, \u0026#34;United States of America\u0026#34;, \u0026#34;United Kingdom\u0026#34;, \u0026#34;France\u0026#34;, \u0026#34;Germany\u0026#34;, \u0026#34;Australia\u0026#34;, \u0026#34;Japan\u0026#34;, \u0026#34;Canada\u0026#34;, \u0026#34;Italy\u0026#34;)) %\u0026gt;% filter(Series == \u0026#34;Emissions per capita (metric tons of carbon dioxide)\u0026#34;) world_plot \u0026lt;- world_data %\u0026gt;% ggplot(mapping = aes( x = factor(Year), y = as.numeric(Value), color = Country, fill = Country, group = Country, )) world_plot + geom_point() + geom_line() + labs( x = NULL, y = \u0026#34;Emissions per capita (metric tons of carbon dioxide)\u0026#34;, title = \u0026#34;Per capita emissions of some countries\u0026#34;, subtitle = \u0026#34;Data from 1975 to 2018\u0026#34;, caption = \u0026#34;Data sources: http://data.un.org/default.aspx\u0026#34;, ) + theme( plot.title = element_text(size = 14), # title position text = element_text(family = \u0026#34;JetBrains Mono\u0026#34;), # font ) Diagram of data by country # Visualize # We just need to add a statement to separate the chart by country on the basis of the above. At the same time, since national data are already included in various charts, we don\u0026rsquo;t need additional legends. world_plot \u0026lt;- world_data %\u0026gt;% ggplot(mapping = aes( x = factor(Year), y = as.numeric(Value), color = Country, fill = Country, group = Country, )) world_plot + geom_point() + geom_line() + facet_wrap(~Country) + labs( x = NULL, y = \u0026#34;Emissions per capita (metric tons of carbon dioxide)\u0026#34;, title = \u0026#34;Per capita emissions of some countries\u0026#34;, subtitle = \u0026#34;Data from 1975 to 2018\u0026#34;, caption = \u0026#34;Data sources: http://data.un.org/default.aspx\u0026#34;, ) + theme( legend.position = \u0026#34;none\u0026#34;, plot.title = element_text(size = 14), # title position text = element_text(family = \u0026#34;JetBrains Mono\u0026#34;), # font ) Conclusions of the study # Canada, France, Germany, Italy and Japan have not changed much overall. The United States has continued to decline in recent years, while China is on the rise. Full code library(tidyverse) library(ggplot2) data_csv \u0026lt;- read.csv(\u0026#34;SYB64_310_202110_Carbon Dioxide Emission Estimates.csv\u0026#34;) world_data \u0026lt;- data_csv %\u0026gt;% filter(Country %in% c(\u0026#34;China\u0026#34;, \u0026#34;United States of America\u0026#34;, \u0026#34;United Kingdom\u0026#34;, \u0026#34;France\u0026#34;, \u0026#34;Germany\u0026#34;, \u0026#34;Australia\u0026#34;, \u0026#34;Japan\u0026#34;, \u0026#34;Canada\u0026#34;, \u0026#34;Italy\u0026#34;)) %\u0026gt;% filter(Series == \u0026#34;Emissions per capita (metric tons of carbon dioxide)\u0026#34;) world_plot \u0026lt;- world_data %\u0026gt;% ggplot(mapping = aes( x = factor(Year), y = as.numeric(Value), color = Country, fill = Country, group = Country, )) world_plot + geom_point() + geom_line() + facet_wrap(~Country) + labs( x = NULL, y = \u0026#34;Emissions per capita (metric tons of carbon dioxide)\u0026#34;, title = \u0026#34;Per capita emissions of some countries\u0026#34;, subtitle = \u0026#34;Data from 1975 to 2018\u0026#34;, caption = \u0026#34;Data sources: http://data.un.org/default.aspx\u0026#34;, ) + theme( legend.position = \u0026#34;none\u0026#34;, plot.title = element_text(size = 14), # title position text = element_text(family = \u0026#34;JetBrains Mono\u0026#34;), # font ) ","date":"19 June 2022","externalUrl":null,"permalink":"/en/posts/carbon-dioxide-emission-estimates/","section":"Blogs","summary":"Data source: Carbon Dioxide Emission Estimates\nIntroduction # The data comes from undata. Last updated on November 1, 2021. It contains the total carbon emissions and per capita carbon emissions of various countries in the world for several years.\n","title":"Carbon Dioxide Emission Estimates","type":"posts"},{"content":"","externalUrl":null,"permalink":"/en/categories/","section":"Categories","summary":"","title":"Categories","type":"categories"},{"content":"The various products that have left a deep impression on me through personal experience, some of which are still in use, while others may have already met their end, or have been resold, or have been discontinued for various reasons, but all of them are recorded here (in reverse order of acquisition), as a cybernetic memorial hall 😄.\n","externalUrl":null,"permalink":"/en/goods/","section":"Goods","summary":"The various products that have left a deep impression on me through personal experience, some of which are still in use, while others may have already met their end, or have been resold, or have been discontinued for various reasons, but all of them are recorded here (in reverse order of acquisition), as a cybernetic memorial hall 😄.\n","title":"Goods","type":"goods"}]