[
  
    {
      "title"       : "On the trend towards the digital twin",
      "category"    : "",
      "tags"        : "digital twin, coding, unreal engine, cesium ion",
      "url"         : "./Toward-The-Digital-Twin.html",
      "date"        : "2023-02-08 13:32:20 +0800",
      "description" : "The trend toward digital twin.",
      "content"     : "From Nature Comments by TaoDigital twins — precise, virtual copies of machines or systems — are revolutionizing industry. Driven by data collected from sensors in real time, these sophisticated computer models mirror almost every facet of a product, process or service. Many major companies already use digital twins to spot problems and increase efficiency1. Half of all corporations might be using them by 2021, one analyst predicts.For instance, NASA uses digital copies to monitor the status of its spacecraft. Energy companies General Electric (GE) and Chevron use them to track the operations of wind turbines. Singapore is developing a digital copy of the entire city to monitor and improve utilities. Machine intelligence and cloud computing will boost such models’ power.There is much to be done to realize the potential of digital twins. Each model is built from scratch: there are no common methods, standards or norms. It can be difficult to aggregate data from thousands of sensors that track vibration, temperature, force, speed and power, for example. And data can be spread among many owners and be held in various formats. For example, the designers of a particular car might hold information on its materials and structure, while the manufacturers keep data on how the vehicle is produced and garages retain information on sales and maintenance.The result? Confusion. A digital twin can fail to echo what is going on in the real world and lead managers to make poor decisions.Here we set out the main problems and call for closer collaboration between industry and academia to solve them.Data difficultiesThe first step is to decide what types of data to collect3. It is not always obvious. To model a wind turbine, for example, might require monitoring of vibrations from the gearbox, generator, blades, shafts and tower, as well as of voltages from the control system. Torques and rotation rates, temperatures of components and the state of the lubricating oil must also be tracked, together with environmental conditions (wind speed, wind direction, temperature, humidity and pressure).Missing or erroneous data can distort results and obscure faults. The wobbling of a wind turbine, say, would be missed if vibration sensors fail. Beijing-based power company BKC Technology struggled to work out that an oil leak was causing a steam turbine to overheat. It turned out that lubricant levels were missing from its digital twin.The optimal number and placing of sensors must be determined. Too few, and predictions will be inaccurate; too many, and the user will be mired in detail. The rate of data collection also matters. Engineers might monitor vibrations from a turbine gearbox every minute, meaning they would miss shorter glitches. But sampling every second could yield way too much data, leading to transmission bottlenecks.To illustrate: Google’s self-driving car could produce 1 gigabyte of data each second, according to some estimates. But today’s bluetooth connections can handle only 0.03% of that rate.Disparate data types are hard to merge, too. Vibrations can be recorded as lengths of time or as frequencies; temperatures can be in celsius or fahrenheit; and videos or images might not be to the same scale. Timings can get out of step, especially when data are sampled at different rates. For example, aircraft communication systems send signals every few nanoseconds, while navigation systems record the position of a plane every second. Averaging fine data doesn’t help because detail is lost.Scattered ownership of data is another barrier. For instance, Boeing aircraft include parts from more than 500 suppliers in 70 countries, each with different data interfaces, formats and software. Companies often don’t want to share commercially sensitive information. Nor do countries: Japan restricts the export of some computer chips to competitors in South Korea, and the United States bans the sale of chips and other technology to Chinese company Huawei.Model challengesTo build a digital twin of an object or system, researchers must model its parts. The German manufacturing company Siemens uses many mathematical models and virtual representations of its products and production lines. These include 3D geometric models and finite-element analysis, the latter for tracking temperatures, stresses and strains. Fault diagnosis and life cycles are treated separately.Other errors could arise when software written for different purposes is patched together by hand. And without standards and guidelines, it is hard to verify the accuracy of the resulting models. Many digital twins might need to be combined. For example, a virtual aircraft might incorporate a 3D model of the fuselage with one of a fault-diagnosis system and one monitoring the air conditioning and pressurization.Even the definition of a digital twin is not settled. Some people think any 3D model or simulation counts. More ambitiously, others envisage a set of integrated models or software that pairs the digital world with physical assets, with or without live information from sensors. Each approach has its own norms, with little crossover.Twin teamsA close-knit team of specialists spanning disciplines is thus essential to building a precise digital twin. No one person can know every detail. Materials scientists, metallurgists and mechanics might need to work with engineers, computer scientists and manufacturing experts. The range of disciplines needed will widen as applications diversify.Most digital twins are found in large companies such as GE or Siemens, because it is difficult and expensive to assemble the teams required. With commercial pressures dissuading businesses from sharing models, smaller firms lose out.There is a lack of common space — physical and virtual — in which experts can communicate and share knowledge and software. And there are few connections between industry and academia, in part because of commercial secrecy. Most academic research focuses on improving modelling techniques rather than on optimizing data and implementing digital twins.Four bridgesThe following steps would make research and development of digital twins more coherent.Unify data and model standards. Manufacturing data should be standardized and delivered in common formats such as XML (Extensible Markup Language), which is used in areas from electronic commerce to supply-chain software. Other data standards should be adopted where they exist. For example, the electricity sector uses COMTRADE (‘common format for transient data exchange’), a standard overseen by the Institute of Electrical and Electronics Engineers; the construction industry uses Industry Foundation Classes; and international health-care organizations require data to conform to HL7 (Health Level 7) standards.A universal design and development platform for digital twins should also be developed on which all models can run. One step in the right direction is a virtual shared workspace, the Global Collaborative Environment, created by aircraft maker Boeing to align practices among its corporate partners. Corporations, foundations, universities and governments should set up and fund an association to oversee a broader one. It could emulate the chip industry’s non-profit research consortium founded in 1982. Called the Semiconductor Research Corporation, this is based in Durham, North Carolina.Share data and models. A public database for sharing digital twins should be created, to be managed by government funding agencies or by a coalition of universities and enterprises. Issues of data ownership and openness will need to be addressed.One such example is the openVertebrate platform, funded by the US National Science Foundation. It allows researchers to freely share data and models of vertebrate anatomies. Digital images and 3D mesh files can be explored, downloaded and 3D-printed on MorphoSource, an open-access online database. Curators can oversee ‘virtual loans’ of their specimen data and receive updates on their use.Platforms of this kind would allow researchers from industry to purchase digital twin data and models, or to lease them to others to conduct research and develop business applications.Innovate on services. Companies should develop products and services to help digital twins become easier to build and use. For example, Siemens’ NX software combines design, simulation and manufacturing tools in one package. Canadian company LlamaZOO has developed a virtual-reality/augmented-reality application that enables mining supervisors to monitor their vehicles. The virtual forest developed by Metsä Group, Tieto and CTRL Reality, all based in Finland, simulates different forest-management methods and their impacts on income and the landscape.Establish forums. Practitioners and researchers need an online space where they can discuss, develop and publish specifications. That is why, in 2017, we set up a social media group on digital twins on the Chinese social-media platform WeChat. Foundations, universities and companies should offer similar forums.Physical ‘innovation hubs’ should also be set up in mutually accessible locations to connect industry, data scientists, cybersecurity experts and engineering and business strategists. One example is the Smart Innovation Hub on the campus of Keele University, UK, alongside Keele Business School. And business consultants Booz Allen Hamilton run several such hubs in Washington DC, near federal government agencies.Nature 573, 490-491 (2019)So we have the comments above. How to think of the next step?Here is a vedio from YouTube https://www.youtube.com/watch?v=UJ3akL3gH68"
    } ,
  
    {
      "title"       : "A brief review for deep Visual Odometry since 2016",
      "category"    : "",
      "tags"        : "machine learning, coding, neural networks",
      "url"         : "./A-brief-review-for-deep-Visual-Odometry.html",
      "date"        : "2022-07-20 19:13:20 +0800",
      "description" : "A systematic overview for Visual Odometry of VSLAM.",
      "content"     : "Application of Deep Learning in Visual Odometry: A Brief Literature ReviewAbstract—Visual odometry is a technique for estimating camera egomotion based on continuous frame images and has important applications in areas such as UAV navigation and augmented reality. Traditional visual odometry mainly applies geometry-based methods, enabling near real-time applications on drones and robots. However, the classical methods have limited applications in challenging cases due to problems such as sensitivity to scene illumination and difficulty in detecting dynamic environments. With the booming development of deep learning in recent years, related techniques combined with visual odometry have emerged as a viable complement, but the current review still focuses on the traditional methods. Therefore, in this paper, we review the classical development of visual odometry to highlight the progress of deep learning incorporating or improving VO traditional methods in recent years and discuss possible current issues and trends.**I. IntroductionVisual odometry is a technique that estimates the camera pose without a priori knowledge by detecting the motion of the surrounding environment. In contrast to VSLAM, which focuses on global consistency, VO focuses mainly on the consistency of local trajectories. The term was coined by D. Nister because vision-based localization is similar to wheel odometry in that it incrementally estimates the motion of a vehicle by integrating the number of turns of its wheels over time[1].Since it was first proposed by D. Nister et al.[2] Since 204, visual odometry has played an important role in many aspects, such as augmented and visual reality, Mars rover exploration, autonomous driving, and navigation of drones. However, these clear methods need another angle of enhancement due to the poor robustness of traditional geometry-based methods in the presence of large differences in illumination and drastic changes in environmental dynamics. With the advancement and development of neural networks and deep learning, the traditional way of using geometric methods in visual mapping has been aided by sophisticated, implicit but more effective deep learning efforts due to their superior performance in vision-related tasks. Considering the cost of training and storing models for these networks, integrating AI-VSLAM with UAVs is a challenge.In this paper, we will review the outstanding contributions of traditional methods in VO and free up more space for an in-depth discussion of highly promising deep learning methods in the VO domain. In contrast to SLAM, in which we are only concerned with the local consistency of trajectories, local maps are used to obtain more accurate estimates of local trajectories (e.g., in bundle adjustment), while SLAM is concerned with the consistency of global maps. This paper is organized as follows. Section II shows some relevant reviews and surveys. Section III provides a brief review of the main previous contributions of DeepVO. Section IV outlines the mainstream ways of combining deep learning and VO. Section V discusses the current potential and challenges of using deep learning as a recognizer at a relatively low cost. The last section summarizes our survey and evaluation.I. Related WorksBecause of the close connection between visual odometry and SLAM technology, VO is often part of SLAM reviews, and there are few reviews specifically on visual odometry compared to SLAM technology. Early classic reviews on VO include a series of tutorials presented by D. Scaramuzza et al[3]. Also, in presenting the development of VO, researchers focused mainly on traditional approaches. Even after 2016, when deep learning has become very popular, reviews and reviews are still scarce, despite the progress made in recent years in visual odometry methods incorporating deep learning.MoShan et al[4]. introduced the application of VIO in MAV; Chen et al[5]. outlined the main applications of VO in SLAM mainly by traditional VO techniques, and also spent a subsection on the integration of deep learning and SLAM in general; in addition, the development of deep learning in SLAM was also discussed in the review by Li et al[6]. However, they mainly introduce the possible directions of deep learning in SLAM such as semantic SLAM, and do not describe the specific potential and methods of deep learning for VO applications; Lai et al[7]. provide a more systematic review of the combination of VSLAM and deep learning, summarize the advantages of the current use of deep learning to deal with SLAM problems, and list the advantages of VSLAM in recent years by applying Li et al[8]. provide a detailed overview of the evolution of VSLAM, but they focus on depth estimation and semantic map building, without a specific collation of VO. Wang et al[9]. give a comprehensive introduction to the methods and applications of deep learning in VO, and list the current problems in the field.I. Geometry-based ApproachesThe traditional implementation methods of VO mainly include the feature point method, direct method, and hybrid semi-direct tracking method.A. Feature Point MethodThe feature point method is a representative method and an early attempt of VO. It can be compared to the principal component of an image, and the feature point method extracts the sparse representative information of the image and estimates the adjacent key frame motion of the whole image based on the overall motion of the feature points. Classical feature point extraction methods include the early Harris corner point[11], FAST corner point[12], and so on. These classical corner point recognition algorithms were proposed earlier and are not stable enough in the case of large image changes. Scale Invariant Feature Transform (SIFT)[13] is a classic algorithm that is robust to illumination, scale, and rotation. However, with the recognition effect comes a huge amount of computation, which makes it difficult for SLAM to meet its real-time requirements. To improve the speed of the algorithm operation, H. Bay et al. proposed Speeded Up Robust Features (SURF)[14] to reduce the computational effort in the SIFT integration process. However, both of these methods still require significant computational costs, and executing the computation in real-time may be challenging.To reconcile accuracy, robustness, and computational effort, Rublee et al[15]. improved the BRIEF descriptor proposed by M. Calonder et al[16]. and addressed the direction invariance of FAST by proposing the oriented FAST and rotated BRIEF (ORB) descriptor. In 2015, based on Klein et al’s Parallel Tracking and Mapping (PTAM)[17], R. Mur-Artal et al.[18] proposed a landmark solution: a robust and accurate real-time ORB-SLAM system. They later improved on this system with ORB-SLAM2[19], which further supported calibrated binocular and RGB-D cameras, and C. Campos et al. went on to introduce ORB3[20], which enriched the sub-maps to improve robustness and further incorporated IMU to enrich the calibration data.B. Direct MethodThe feature point method is clear and straightforward, but there are still some problems. For example, even if the ORB speed is already quite fast, it still takes about 20ms[21], and if we want to do a 30-frame real-time SLAM, then we need each frame to be around 33ms on average. Thus, most of the time overhead is spent on feature point extraction. In addition, the image itself is discarded when feature points are used. Although feature points can reflect the image in a sense, an image has after all millions of pixels, and feature points are often only a few hundred, which may be difficult to reflect potentially useful image information in some cases. Meanwhile, in some occasions where feature points are not too significant, such as along the direction of a wall or an empty corridor, it is difficult to identify the camera movement by feature points alone. All these may pose related problems. That is, there is a certain relationship among them.Therefore, in some cases, the direct method may be more appropriate. In contrast to the feature point method, the direct method does not require a one-to-one match, and the projection is considered successful as long as the previous points have reasonable projection residuals in the current image: success depends mainly on the judgment of the depth of the map points and the camera pose, not on what the image looks like locally. The direct method saves a lot of time in feature extraction and matching is easily portable to embedded systems and can be integrated with IMUs, of which the LK optical flow technique[22] is a well-known approximate example. Since its introduction, the optical flow method has been continuously developed[23]. Direct methods like this seem to directly use image pixel grayscale information and geometric information to construct error functions by graphical optimization to minimize the cost function and thus obtain the best camera pose. In practice, Engel et al. proposed the large-scale direct monocular simultaneous localization and mapping (LSD-SLAM) algorithm[24] and applied it to a stereo camera, combining temporal and static stereo in a direct, real-time SLAM approach[25]. realistic conditions with some robustness considering also illumination variations. After this, Engel et al. further proposed DSO-SLAM[26]. however, its parameters in the code need to be adjusted to adapt to the new scene requirements each time the scene is changed in a dynamic environment, and there are problems such as scale drift in practical application scenarios.Deep Learning ApproachesHowever, although geometry-based SLAM has been able to achieve CPU real-time in classical scenes, traditional, geometry-based SLAM methods still have some problems: for the feature point method, identifying feature points may encounter some difficulties in the case of insignificant features. In addition, additional arithmetic power is required to extract features, and these computations account for most of the entire VO process, and these feature points are discarded soon after matching, resulting in a large degree of waste; for the direct optical flow method, the assumption of constant features such as overall illumination of the rigid scene is required, and these are difficult to implement in scenes such as outdoors. Therefore, with the development of deep learning and its great advantages shown in visual recognition, many VOs incorporating deep learning have been proposed. Convolutional neural networks, a network structure, were the first to come into view due to their dominant level in object recognition and detection problems.Review of Supervised Deep LearningIn this context, Kendall et al[27]. proposed PoseNet capable of generating six degrees of freedom of a camera directly from a single RGB input image and was the first implementation of camera pose estimation. Since CNN extracts more powerful features than conventional feature detectors, the system can achieve high accuracy even under certain extreme conditions, such as strong illumination and blurred images. Later the authors improved PoseNet and also proposed improvements based on Bayesian analysis[28], which improved the accuracy of relocation; another direction of improvement by the authors was to improve the performance of PoseNet by using multi-view geometry as a source of training data[29]. Li et al[30]. extended PoseNet to accommodate color and depth inputs from RGB-D cameras using a dual-stream convolutional neural network, which showed robust performance in the face of challenging situations, becoming the first work to solve the deep CNN-based indoor relocation problem using RGB-D cameras. Wang et al[31]. proposed DeepVO, a recurrent convolutional neural network (RCNN)-based VO approach that is competitive with model-based VO approaches, as a notable advance.Review of unsupervised or self-supervised Deep LearningAll these above are the applications of supervised learning methods in VO: supervised learning methods tend to obtain better pose estimation results. However, SLAM is a niche area where it is often difficult and expensive to obtain real ground truth datasets in practice. It is difficult to build datasets suitable for large supervised learning and to label ground truth, while the number of available labeled datasets for supervised training is still limited. Li et al[32]. proposed a new monocular visual ranging (VO) system called UnDeepVO, which can achieve recovery of absolute scale. As an unsupervised approach, compared to DeepVO, UnDeepVO can be trained using a large number of unlabeled datasets to continuously improve its performance. Ummenhofer proposed DeMoN[33], which can simultaneously estimate camera self-motion, image depth, surface normals, and optical flow. Compared to popular single-image depth networks, DeMoN learns the concept of matching and thus can be better generalized to structures not seen during training. both DeMoN and UnDeepVO use stereo images to train the network to eliminate the important scale ambiguity problem in monocular V and are the first network models to use unsupervised learning methods to estimate the depth and pose of continuous images. The GANVO[34] proposed by Almalioglu et al. In contrast to traditional VO methods, pose and depth estimation does not require strict parameter tuning while being able to address the problem that traditional depth estimators based on autoencoder decoders tend to generate overly smooth images. However, unsupervised methods suffer from the drawback of insufficient supervisory information, so self-supervised algorithms by adding known image features as supervisory signals are also widely proposed. The D3VO self-supervised monocular depth estimation network proposed by Yang et al. tightly combines predicted depth, pose, and uncertainty into a direct visual ranging approach to enhance front-end tracking and back-end nonlinear optimization .e k can be analyzability, etc.Methods ComparisonAs shown in Table 1, due to the rapid development of deep learning in VO applications in recent years, this paper collates the progress of the main network models according to the characteristics of different models. The collation criteria include five main dimensions to evaluate: the use of network structure, the type of training (supervised or not), the test dataset used, whether it is an end-to-end model, and the type of camera applied to the model.TABLE I. Characteristics of the major deep learning Visual Odometry models Name Structure Type Benchmark E2E Sensor PoseNet CNN Supervised Cambridge Landmarks, 7 Scenes dataset[35] Y Mono DeepVO RCNN Supervised KITTI[36] Benchmark Y Mono D3VO DeepThingNet Self-supervised KITTI &amp; EuRoC[37] N Mono UndeepVO RCNN Unsupervised KITTI Benchmark Y Mono GANVO GAN Unsupervised KITTI &amp; Cityscapes[38] Y Mono DeMoN Bootstrap Net Supervised SUN3D[39] &amp; MVS[40] Y Mono From the above table for the summary of influential models in recent years, it can be seen that due to the complexity and specificity of VO, the applied network structure has changed from the early CNN ruling the situation to the present blossoming; in addition, the number of available large-scale datasets in VO or SLAM is still limited due to the development of VO datasets suitable for deep learning slightly lagging behind the development of network models.As a result, unsupervised and self-supervised approaches have been emerging since 2017 and can obtain stronger generalization while maintaining accuracy. In addition, since neural networks can be likened to a black box, most applications have adopted an end-to-end model, i.e., replacing the process from feature extraction to camera pose estimation in traditional VO methods; in terms of the cameras used, thanks to the advantage of deep learning in reducing the estimated absolute depth, researchers in most application scenarios favor monocular cameras to reduce the cost and improve the generalization capability, researchers have favored monocular cameras in most applications to reduce costs and improve generalization capabilities.Further DiscussionsSince 2015, deep learning has been increasingly integrated with VO technology. Rich applications have sprung up based on traditional feature extraction or optical flow computation. However, the current research still has some possible problems. The next section will discuss the challenges faced and possible directions for development.Challenges and DifficultiesDynamic scenes and dynamic objects in the scenes. VO assumes that the environment is static to integrate the egomotion of the camera, however, the scenes in which VO is performed are likely to encounter dynamic objects, such as pedestrians, animals, etc.; in this regard, the illumination may change more drastically, for example, the illumination may not be uniformly distributed in outdoor environments.Deep learning is poorly interpretable, and because the training set cannot contain all scenes, the visual odometry trained by deep learning is often limited to certain specific scenes and performs poorly in some unfamiliar scenes.Insufficient training data. In addition to the KITTI benchmark and Cityscapes datasets mentioned above, the mainstream VO datasets available include the RobotCar dataset[41], which contains different weather and scenery in the same location and was collected using a car driven in Oxford for a year. Also available is the previously mentioned EuRoc MAV dataset, which is a dataset that can be used for VO and VSLAM by collecting data through MAV. All of these datasets can be used for self-motion estimation, and the Cityscapes and KITTI datasets can also be used to complete scene segmentation. Although these training sets are relatively rich, they tend to be limited to certain scenes, for which overfitting may lead to a decrease in the model’s generalization ability and struggle to perform in some unfamiliar environments, which is exactly where VO tends to run.####Prospects and DirectionsUse semantic predictive feedback to reduce the interference of dynamic objects in the scene. Semantic segmentation of the scene is performed during image processing and the results of the semantic segmentation are used as a correction to modify the operation of the VO. For example, Barnes et al[42]. distinguish static and dynamic parts of the scene by integrating a per-pixel ad-hoc mask in the VO to determine unreliable regions in the image.Unsupervised learning does not require much hard-to-label ground truth, and thus semi-supervised, self-supervised, or unsupervised methods can be used to reduce the requirement for training set labeling when the dataset is not fully developed.Fuse additional information such as IMU into the network structure for loss processing, while giving each other synchronization feedback, that is, in the direction of Deep VIO.​ For the problem of poor model prediction improvement due to poor interpretability of deep learning, the degree of overfitting can be reduced by methods in artificial intelligence. Yang et al[43]. introduce the Bayesian distribution of weight factors to improve the generalization ability of network models in the prediction process and improve certain robustness for translation and rotation; they can also provide improved network structures to enhance the robustness for scenarios that do not appear in the training set and generalization.ConclusionsIn this paper, we review the classical methods of VO, and on this basis, we make a brief arrangement and summary of the fusion and application of deep learning in VO in recent years. From the summary, we can see that compared with the traditional methods, the deep learning methods have good results in the case of very sparse or insignificant features. As deep learning continues to develop recognition capabilities in various visual tasks, research attention is increasingly turning toward deep learning and VO fusion. Despite the current shortcomings compared to clear solutions with geometry, they have shown great potential in various areas of VO applications.AcknowledgmentThis work was inspired by A brief survey of visual odometry for micro aerial vehicles and other prominent works by professor *Ben M. Chen*. Appreciation for the inspiration and guidance from his articles that motivated me to manage the research.References[1] Scaramuzza D, Fraundorfer F, Pollefeys M. Closing the loop in appearance-guided omnidirectional visual odometry by using vocabulary trees. Robot Auton Syst. 2010;58(6):820–827. doi: 10.1016/j.robot.2010.02.013.[2] D. Nister， O. Naroditsky, and J. Bergen， “Visual odometry”， Proc. Int. Conf. Computer Vision and Pattern Recognition， pp. 652-659， 2004.[3] D. Scaramuzza and F. Fraundorfer, “Visual odometry. part i: The rst 30 years and fundamentals”, IEEE Robot. Autom. Mag, vol. 18, pp. 8092, 2011.[4] Mo Shan et al., “A brief survey of visual odometry for micro aerial vehicles,” IECON 2016 - 42nd Annual Conference of the IEEE Industrial Electronics Society, 2016, pp. 6049-6054, doi: 10.1109/IECON.2016.7793198.[5] Y. Chen， Y. Zhou， Q. Lv and K. K. Deveerasetty， “A Review of V-SLAM”， 2018 IEEE International Conference on Information and Automation （ICIA）， 2018， pp. 603-608， doi： 10.1109/ICInfA.2018.8812387.[6] A. Li， X. Ruan， J. Huang， X. Zhu and F. Wang， “Review of vision-based simultaneous Localization and Mapping，” 2019 IEEE 3rd Information Technology， Networking， Electronic and Automation Control Conference （ITNEC）， 2019， pp. 117-123， doi： 10.1109/ITNEC.2019.8729285.[7] D. Lai， Y. Zhang 和 C. Li， “A Survey of Deep Learning Application in Dynamic Visual SLAM”， 2020 International Conference on Big Data &amp; Artificial Intelligence &amp; Software Engineering （ICBASE）， 2020， pp. 279-283， doi： 10.1109/ICBASE51474.2020.00065.[8] Li, R., Wang, S. &amp; Gu, D. Ongoing Evolution of Visual SLAM from Geometry to Deep Learning: Challenges and Opportunities. Cogn Comput 10, 875–889 (2018). https://doi.org/10.1007/s12559-018-9591-8[9] Wang, S. Ma, J. Chen, F. Ren and J. Lu, “Approaches, Challenges, and Applications for Deep Visual Odometry: Toward Complicated and Emerging Areas,” in IEEE Transactions on Cognitive and Developmental Systems, vol. 14, no. 1, pp. 35-49, March 2022, doi: 10.1109/TCDS.2020.3038898.[10] He, M., Zhu, C., Huang, Q. et al. A review of monocular visual odometry. Vis Comput 36, 1053–1065 (2020). https://doi.org/10.1007/s00371-019-01714-6[11] C. Harris and M. Stephens, “A combined corner and edge detector”, Proc. Alvey Vis. Conf., vol. 15, no. 50, pp. 5244, 1988.[12] Rosten, E., Drummond, T.: Machine learning for high-speed corner detection. In: European Conference on Computer Vision, pp. 430–443. Springer, Berlin (2006)[13] D. G. Lowe, “Distinctive image features from scale-invariant keypoints”, International Journal of Computer Vision (IJCV), vol. 60, no. 2, pp. 91-110, 2004.[14] H. Bay, A. Ess, T. Tuytelaars and L. V. Gool, “SURF: Speeded up robust features”, Computer Vision and Image Understanding (CVIU), vol. 110, no. 3, pp. 346-359, 2008.[15] Rublee, E., Rabaud, V., Konolige, K., Bradski, G.: ORB: an efficient alternative to SIFT or SURF. In: 2011 IEEE international conference on computer vision (ICCV), pp. 2564–2571. IEEE (2011)[16] M. Calonder, V. Lepetit, C. Strecha, and P. Fua. Brief: Binary robust independent elementary features. In In European Conference on Computer Vision, 2010. 1, 2, 3, 5[17] Klein, G., Murray, D.: Parallel tracking and mapping for small AR workspaces. In: 6th IEEE and ACM International Symposium on Mixed and Augmented Reality, 2007 (ISMAR 2007), pp. 225–234. IEEE (2007)[18] R. Mur-Artal, J. M. M. Montiel and J. D. Tardós, “ORB-SLAM: A Versatile and Accurate Monocular SLAM System,” in IEEE Transactions on Robotics, vol. 31, no. 5, pp. 1147-1163, Oct. 2015, doi: 10.1109/TRO.2015.2463671.[19] R. Mur-Artal and J. D. Tardós, “ORB-SLAM2: An Open-Source SLAM System for Monocular, Stereo, and RGB-D Cameras,” in IEEE Transactions on Robotics, vol. 33, no. 5, pp. 1255-1262, Oct. 2017, doi: 10.1109/TRO.2017.2705103.[20] C. Campos, R. Elvira, J. J. G. Rodríguez, J. M. M. Montiel and J. D. Tardós, “ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual–Inertial, and Multimap SLAM,” in IEEE Transactions on Robotics, vol. 37, no. 6, pp. 1874-1890, Dec. 2021, doi: 10.1109/TRO.2021.3075644.[21] F. Endres, J. Hess, N. Engelhard, J. Sturm, D. Cremers and W. Burgard, “An evaluation of the RGB-D SLAM system,” 2012 IEEE International Conference on Robotics and Automation, 2012, pp. 1691-1696, doi: 10.1109/ICRA.2012.6225199.[22] B.D. Lucas and T. Kanade, “An iterative image registration technique with an application to stereo vision”, Proc. DARPA Image Understanding Workshop, pp. 121-130, 1981.[23] Baker, S., Matthews, I.: Lucas-Kanade 20 years on: a unifying framework. Int. J. Comput. Vis. 56(3), 221–255 (2004)[24] Engel, J., Schöps, T., Cremers, D.: LSD-SLAM: large-scale direct monocular SLAM. In: European Conference on Computer Vision, pp. 834–849. Springer, Cham (2014)[25] J. Engel, J. Stückler and D. Cremers, “Large-scale direct SLAM with stereo cameras,” 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2015, pp. 1935-1942, doi: 10.1109/IROS.2015.7353631.[26] J Engel, V Koltun, D Cremers et al., “Direct sparse odometry[J]”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 3, pp. 611, 2018.[27] A. Kendall, M. Grimes and R. Cipolla, “PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization,” 2015 IEEE International Conference on Computer Vision (ICCV), 2015, pp. 2938-2946, doi: 10.1109/ICCV.2015.336.[28] A. Kendall and R. Cipolla, “Modelling uncertainty in deep learning for camera relocalization”, Proc. IEEE Int. Conf. Robot. Autom. (ICRA), pp. 4762-4769, May 2016.[29] A. Kendall and R. Cipolla, “Geometric loss functions for camera pose regression with deep learning”, Proc. 30th IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 6555-6564, 2017.[30] R. Li, Q. Liu, J. Gui, D. Gu and H. Hu, “Indoor relocalization in challenging environments with dual-stream convolutional neural networks”, IEEE Transactions on Automation Science and Engineering, 2017.[31] S. Wang， R. Clark， H. Wen 和 N. Trigoni， “DeepVO： Towards end-to-end visual odometertry with deep recurrent convolutional neural networks”， Robotics and Automation （ICRA） 2017 IEEE International Conference on， pp. 2043-2050， 2017.[32] R. Li, S. Wang, Z. Long and D. Gu, “UnDeepVO: Monocular Visual Odometry Through Unsupervised Deep Learning,” 2018 IEEE International Conference on Robotics and Automation (ICRA), 2018, pp. 7286-7291, doi: 10.1109/ICRA.2018.8461251.[33] B. Ummenhofer， H. Zhou， J. Uhrig， N. Mayer， E. Ilg， A. Dosovitskiy， et al.， “DeMoN： Depth and Motion Network for learning monocular stereo”， Conference on Computer Vision and Pattern Recognition （CVPR）， 2017.[34] Y. Almalioglu, M. R. U. Saputra, P. P. B. d. Gusmão, A. Markham and N. Trigoni, “GANVO: Unsupervised Deep Monocular Visual Odometry and Depth Estimation with Generative Adversarial Networks,” 2019 International Conference on Robotics and Automation (ICRA), 2019, pp. 5474-5480, doi: 10.1109/ICRA.2019.8793512.[35] J. Shotton, B. Glocker, C. Zach, S. Izadi, A. Criminisi and A. Fitzgibbon, “Scene coordinate regression forests for camera relocalization in RGB-D images”, Computer Vision and Pattern Recognition (CVPR) 2013 IEEE Conference on, pp. 2930-2937, 2013.[36] Andreas Geiger, Philip Lenz and Raquel Urtasun, “Are we ready for autonomous driving? the KITTI vision benchmark suite”, Conference on Computer Vision and Pattern Recognition (CVPR), 2012.[37] Michael Burri, Janosch Nikolic, Pascal Gohl, Thomas Schneider, Joern Rehder, Sammy Omari, et al., “The EuRoC micro aerial vehicle datasets”, The International Journal of Robotics Research, 2016.[38] M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, et al., “The cityscapes dataset for semantic urban scene understanding”, Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3213-3223, 2016.[39] J. Xiao， A. Owens and A. Torralba， “SUN3D： A Database of Big Spaces Reconstructed Using SfM and Object Labels”， IEEE International Conference on Computer Vision （ICCV），pp. 1625-1632， Dec. 2013.[40] S. Fuhrmann， F. Langguth and M. Goesele， “Mve-a multiview reconstruction environment”， Proceedings of the Eurographics Workshop on Graphics and Cultural Heritage （GCH）， vol. 6， pp. 8， 2014.[41] Maddern W, Pascoe G, Linegar C, Newman P. 1 Year, 1000km: the Oxford robotCar dataset. The International Journal of Robotics Research (IJRR) 2017;36(1):3–15.[42] D. Barnes, W. Maddern, G. Pascoe and I. Posner, “Driven to distraction: Self-supervised distractor learning for robust monocular visual odometry in urban environments”, Proc. IEEE Int. Conf. Robot. Autom. (ICRA), pp. 1894-1900, 2018.[43] X. Yang, X. Li, Y. Guan, J. Song and R. Wang, “Overfitting reduction of pose estimation for deep learning visual odometry,” in China Communications, vol. 17, no. 6, pp. 196-210, June 2020, doi: 10.23919/JCC.2020.06.016."
    } ,
  
    {
      "title"       : "A systematic overview for Visual Odometry of VSLAM",
      "category"    : "",
      "tags"        : "machine learning, coding, neural networks",
      "url"         : "./A-systematic-overview-for-Visual-Odometry-of-VSLAM.html",
      "date"        : "2022-06-30 19:13:20 +0800",
      "description" : "A systematic overview for Visual Odometry of VSLAM.",
      "content"     : "Visual Odometry技术 （Of VSLAM)[toc]什么是SLAM​ SLAM是Simultaneous localization and mapping缩写，意为“同步定位与建图”1。它是指搭载了特定传感器的主体，如机器人或者无人机等，在没有关于环境的先验知识的情况下，在运动的过程中建立环境的模型。SLAM的概念早在1986年2就提出了。然而，早期的SLAM往往依赖价格昂贵或专门定制的传感器，例如激光雷达，声呐或立体相机，这项技术并未走入市场。随着算法和算力的不断发展，廉价的相机逐渐成为一种取代激光雷达等复杂设备进行SLAM的可能 。那么，如果涉及到的传感器主要为相机，那么就称为“视觉(Visual)SLAM”， 也就是标题所叙述的、这篇博客主要探讨的VSLAM。经典视觉SLAM框架如图1所示，下面是经典视觉SLAM框架的主要组成结构：一个SLAM系统主要由Visual Odometry(视觉里程计，VO)，Optimization(后端优化),Loop Closing(回环检测), Mapping(建图)这四个部分组成。 图1 经典视觉SLAM框架 其中，VO能够通过相邻帧间的图像估计相机运动，并回复场景的空间结构。然而，仅仅有VO是不够的。VO其实就有点像马尔科夫链（Markov Chain）那样，只关注当前状态和未来状态的基本联系，不具有记忆性。这样一来，VO由于只有🐟的记忆，而每次的估计又会有一定偏差，每次估计的相机位姿运动偏差在机器人或者无人机运动过程中不断累加，形成累计偏移（Accumulating Drift）这些累计偏移有的时候会带来极为糟糕的后果。如图2所示，设想一下，如果在估计的时候认为相机顺时针运动了90度，而实际上相机仅仅运动了89度，这样一来在一个空的矩形的房间里所作出的定位可能会因此不断远离相机的实际位置，而建立出来的地图也很可能会无法封闭。 图2 逐渐增大的偏移和无法闭合的地图 因此，我们需要在一个更宏观的视角下审视并且修正这些偏移。这样就引入了后端优化和回环检测这两个部分。在后端优化中，则需要考虑相对更加长远的目标：在解决“如何从图像估计相机运动”的基础上从带有噪声的数据中估计整个系统的状态，以及这个状态估计的不确定性有多大，同时使得得到的相机位姿在全局上尽可能保持一致。相比于VO部分，Optimization往往没有那么可见，面对的只有数据，而不必关心数据的来自于激光雷达还是单目相机，又在视觉里程计之后，因此叫做后端。但是还是要注意的是，这里的前后端和web应用（如J2EE中的Servlet）中的前后端有所不同，要分清楚两者之间的区别。在回环检测部分，最主要判断的还是机器人或者无人机是否达到之前到过的位置，如果检测到了闭环（往往是通过图像相似度判断实现），就会把信息提供给后端进行处理来得到一个全局一致的估计。用一个不太准确但是我自己觉得非常形象的比喻来说，这个过程就像是用把一个个用小棍子（预测的轨迹）穿起来的珠子（估计的点）头尾相接到一起保持中间各个珠子距离不变一样。现阶段应用最广的回环检测方法是词袋模型（Bag-of-Words），之后会详细介绍。下面将对VSLAM框架中的VO技术进行具体介绍。Visual Odometry如上所述，Visual Odometry主要是计算图像帧之间 的相机位姿关系，也即通过拍摄图像，估计出相机的运行位置和姿态信息。根据所使用相机的类型，我们可以把VO分为单目VO和立体VO3，其中单目VO主要使用单目相机来获取环境的2D信息；而立体VO如RGB-D相机和双目相机在获取画面外能够直接通过结构光或者ToF获取场景深度信息或通过计算获得的场景深度信息（类似人眼）在实际操作中，由于RGB-D相机由于很容易受到自然光的干扰，同时对于噪声的鲁棒性较差，本身价格也比较高不利于推广，因此主要用于室内SLAM；而双目相机的精度和深度方向上的量程受到基线长度，也即两个相机间距离的影响（但是做的宽一个是容易形变导致误差，一个是相机太宽影响运动），同时disparity map的计算要消耗大量的资源，往往需要GPU或者FPGA来加速，在深度上的测量很难达到令人满意的效果。因此，单目相机SLAM技术便是这篇博客所要探讨的主要内容。在实践中，VO算法主要分为特征点法和直接法两类。特征点法是通过汇总图像中所有有代表性的点的移动来预测相机的整体移动情况。由于通过矩阵在整个图像的层面来判断运动是十分困难的（LK光流需要强假设），因此我们可以用另一种图像的表现形式，也就是图像的特征来描述图片，减少不必要的信息（特征也可以看作是图像的主成分）。尽管特征点在面对墙体或者其他角点不显著（salient）的区域时可能难以识别4，但是在绝大多数场景下都能够找到充足的特征点来对帧间运动做出一个大致的估计。传统的寻找特征点的方法主要包括Harris角点（参考BUAASE_CV_hw_set2）、FAST角点5等，这些经典的角点识别算法提出的时间较早，在图像变化幅度较大的情况下不够稳定。近年来不断发展的局部特征识别往往不仅匹配角点（或者也可以说兴趣点）本身，还会为角点提供相应的描述子（descriptor）来说明特征点的朝向和大小等信息。例如，SIFT就是一个十分经典的算法，能够对关照、尺度以及旋转都有很好的鲁棒性。然而，随着识别效果而来的还有巨大的计算量。与SfM不同，SLAM要求实时性，因此在课上熟知的SIFT很少被应用到SLAM的实际应用中。那么有没有什么能够协调好准确率、鲁棒性以及计算量，使之能够适配SLAM的算法呢？当然有！这就是在SLAM中大名鼎鼎的ORB-SLAM，如图所示，就像YOLO一样，ORB也更新了很多版，证明了其强大的生命力。 图3 orb各个版本的论文（图源本人） ORB（Oriented FAST and Rotated BRIEF）,是目前最快速稳定的特征点检测和提取算法，许多图像拼接和目标追踪技术利用ORB特征进行实现6。ORB-SLAM 是西班牙 Zaragoza 大学的 Raúl Mur-Arta 编写的视觉 SLAM 系统，目前开源在raulmur/ORB_SLAM: A Versatile and Accurate Monocular SLAM (github.com)上。正如GitHub上的md所述， ORB是一个通用高效的单目SLAM解决方案（后面的ORB2、ORB3支持了更多相机，但是这里先讨论单目的ORB）。就像刚刚提到的，FAST角点以速度快而著称。已FAST-9为例，在像素点为中心的一个半径等于3像素的离散化的Bresenham圆找9个连续的像素点，如果它们们的像素值要么都比中心点加上一个阈值t大，要么都比中心点减去一个阈值t小，那么这个点就是一个角点。注意到FAST只用到了一个圆而没有具体的方向，事实上对于旋转缺乏鲁棒性；此外，由于它固定取半径为3的圆，因此存在尺度问题：可能一些在远处看是角点的位置放大后周围像素趋于一致而不再是角点了。ORB对FAST的改进或者拓展，主要是为其增加了其尺度不变性以及旋转不变性。尺度不变性主要是通过图像金字塔例如Gaussian pyramid向下采样（使用高斯核对其进行卷积，然后对卷积后的图像进行下采样，反复迭代），是一种从粗糙到不断精细的过程。图像金字塔是单个图像的多尺度表示法，由一系列原始图像的不同分辨率版本组成。金字塔的每个级别都由上个级别的图像下采样版本组成。下采样是指图像分辨率被降低，比如图像按照 1/2 比例下采样。因此一开始的 4x4 正方形区域现在变成 2x2 正方形。图像的下采样包含更少的像素，并且以 1/2 的比例降低大小。这样一来，上面提到的角点放大丢失的问题就能够得到解决。 图4 Gaussian pyramid 降采样 旋转不变性主要依靠ORB_SLAM的灰度质心法来处理：灰度质心法首先要选择某个图像块B然后将图像块B的矩m定义为那么可以找到图像块B的质心C:方向向量可以通过将图像块的几何中心和它的质心连接在一起得到，所以可以定义特征点的方向为：这样一来，FAST就有了尺度和方向的描述，就成为了Oriented FAST。那么，有了关键点以后，我们就需要对每个点计算描述子。ORB中使用的描述子是Rotated BRIEF，是BRIEF的一种改进算法。BRIEF 是 Binary Robust Independent Elementary Features 的简称，它的作用是根据一组关键点创建二进制特征向量，又称为二进制特征描述符，是仅包含 1 和 0 的特征向量。在 BRIEF 中 每个关键点由一个二进制特征向量描述，该向量一般为 128-512 位的字符串，其中仅包含 1 和 0。由于向量中的每个值都是0或者1，因此BRIEF很适合用在像SLAM这样对实时性要求高同时资源又特别受限的技术上。BRIEF流程简单实时性较好，论文中生成512个描述子用时8.18ms6，并且其描述子是二进制码，其匹配也比较快。但是，当BRIEF对于旋转过大时，比如超过30度时，匹配正确率迅速下降直到45度时为0。因此，和之前提到的FAST一样，BRIEF也不满足图像的尺度旋转不变性，因此，为使特征满足尺度不变性， Rotated BRIEF 算法同样构建图像金字塔。值得一提的是论文中提到的steered BRIEF 来增加其旋转不变性6：所谓steered BRIEF就是对挑选出的点对加上一个旋转角度θ。对于任何一个特征点来说，它的BRIEF描述子是一个长度为𝑛的二值码串，这个二值码串是由特征点邻域𝑛个点对生成的。在代码实现的过程中我们可以使用OpenCV的库函数来辅助进行Oriented FAST的检测和BRIEF 描述子的计算。//-- 第一步:检测 Oriented FAST 角点位置 chrono::steady_clock::time_point t1 = chrono::steady_clock::now(); detector-&gt;detect(img_1, keypoints_1); detector-&gt;detect(img_2, keypoints_2);//-- 第二步:根据角点位置计算 BRIEF 描述子descriptor-&gt;compute(img_1, keypoints_1, descriptors_1); descriptor-&gt;compute(img_2, keypoints_2, descriptors_2); chrono::steady_clock::time_point t2 = chrono::steady_clock::now(); chrono::duration&lt;double&gt; time_used = chrono::duration_cast&lt;chrono::duration&lt;double&gt;&gt;(t2 - t1); cout &lt;&lt; \"extract ORB cost = \" &lt;&lt; time_used.count() &lt;&lt; \" seconds. \" &lt;&lt; endl; Mat outimg1; drawKeypoints(img_1, keypoints_1, outimg1, Scalar::all(-1), DrawMatchesFlags::DEFAULT); imshow(\"ORB features\", outimg1);//当然，也可以自己实现一个版本void ComputeORB(const cv::Mat &amp;img, vector&lt;cv::KeyPoint&gt; &amp;keypoints, vector&lt;DescType&gt; &amp;descriptors) { const int half_patch_size = 8; const int half_boundary = 16; int bad_points = 0; for (auto &amp;kp: keypoints) { if (kp.pt.x &lt; half_boundary || kp.pt.y &lt; half_boundary || kp.pt.x &gt;= img.cols - half_boundary || kp.pt.y &gt;= img.rows - half_boundary) { // outside bad_points++; descriptors.push_back({}); continue; } float m01 = 0, m10 = 0; for (int dx = -half_patch_size; dx &lt; half_patch_size; ++dx) { for (int dy = -half_patch_size; dy &lt; half_patch_size; ++dy) { uchar pixel = img.at&lt;uchar&gt;(kp.pt.y + dy, kp.pt.x + dx); m10 += dx * pixel; m01 += dy * pixel; } } // angle should be arc tan(m01/m10); float m_sqrt = sqrt(m01 * m01 + m10 * m10) + 1e-18; // avoid divide by zero float sin_theta = m01 / m_sqrt; float cos_theta = m10 / m_sqrt; // compute the angle of this point DescType desc(8, 0); for (int i = 0; i &lt; 8; i++) { uint32_t d = 0; for (int k = 0; k &lt; 32; k++) { int idx_pq = i * 32 + k; cv::Point2f p(ORB_pattern[idx_pq * 4], ORB_pattern[idx_pq * 4 + 1]); cv::Point2f q(ORB_pattern[idx_pq * 4 + 2], ORB_pattern[idx_pq * 4 + 3]); // rotate with theta cv::Point2f pp = cv::Point2f(cos_theta * p.x - sin_theta * p.y, sin_theta * p.x + cos_theta * p.y) + kp.pt; cv::Point2f qq = cv::Point2f(cos_theta * q.x - sin_theta * q.y, sin_theta * q.x + cos_theta * q.y) + kp.pt; if (img.at&lt;uchar&gt;(pp.y, pp.x) &lt; img.at&lt;uchar&gt;(qq.y, qq.x)) { d |= 1 &lt;&lt; k; } } desc[i] = d; } descriptors.push_back(desc); } cout &lt;&lt; \"bad/total: \" &lt;&lt; bad_points &lt;&lt; \"/\" &lt;&lt; keypoints.size() &lt;&lt; endl;}在取得了所有匹配好的点对后，就可以通过点对之间的关系来估计单目相机的运动，这部分由于涉及到大量的对极几何约束而且相关的代码都可以在OpenCV中找到，比如cvFindFundamentalMat和findEssentialMat等函数可以直接免去大量的理解，主要还是理清基本的三角测量原理和本质矩阵的应用，这里就不再过多展开了。 图5 三角测量原理和对极约束 上述便是特征点法操作的主要流程。特征点法很清晰直接，但是还是有一些问题。例如，即使ORB速度已经相当快了，也仍然需要20ms的时长，如果想要做到一个30帧的实时SLAM，那么就需要每一帧在平均33ms左右。这样一来大部分的时间开销都会花在特征点的提取上。此外，使用特征点后，图像本身就被丢弃了。尽管特征点能够在某种意义上反映图像的情况，但是一张图像毕竟有十几万像素，而特征点往往只有几百个，在一些情况下可能难以反映可能有用的图像信息。同时，在一些特征点不是太显著的场合，比如沿着墙体的方向，或者是空无一物的走廊，单靠特征点很难识别出相机的运动。这些都有可能会带来相关的问题。因此，在一些情况下，直接法可能更加使用，其中一个主要的代表就是LSD-SLAM。相比于特征点法，直接法并不要求一一对应的匹配，只要先前的点在当前图像当中具有合理的投影残差，就认为这次投影是成功的：成功与否主要取决于对地图点深度以及相机位姿的判断，并不在于图像局部看起来是什么样子。直接法节省特征提取与匹配的大量时间，易于移植到嵌入式系统中，以及与IMU进行融合，其中LK光流技术就是一个著名的近似例子。Lucas–Kanade光流Lucas–Kanade光流算法是一种两帧差分的光流估计算法。它由Bruce D. Lucas 和 Takeo Kanade提出7。Optical flow, 或者说光流，是一种运动模式，这种运动模式指的是一个物体、表面、边缘在一个视角下由一个观察者（比如眼睛、摄像头等）和背景之间形成的明显移动。它计算两帧在时间t 到t + δt之间每个每个像素点位置的移动。 由于它是基于图像信号的泰勒级数，这种方法称为差分，这就是对于空间和时间坐标使用偏导数。 图像约束方程可以写为\\(I (x ,y ,z,t ) = I (x + δx ,y + δy ,z + δz ,t + δt )\\)其中，I(x, y,z, t) 为在（x,y,z）位置的体素。 我们假设移动足够的小，那么对图像约束方程使用泰勒公式，我们可以得到：忽略高阶无穷小项，可以得到：利用滑动窗口可以得到一个超定方程，使用最小二乘法求解就可以得到：这也是进行估计的基础。此外，LK光流算法也可以通过前面提到的金字塔方法来进行优化，这样一来能够避免运动速度过快、图像整体结构发生变化等问题。在实现中，我们也可以使用高斯牛顿法来计算光流。核心函数如下（cpp）void OpticalFlowTracker::calculateOpticalFlow(const Range &amp;range) { // parameters int half_patch_size = 4; int iterations = 10; for (size_t i = range.start; i &lt; range.end; i++) { auto kp = kp1[i]; double dx = 0, dy = 0; // dx,dy need to be estimated if (has_initial) { dx = kp2[i].pt.x - kp.pt.x; dy = kp2[i].pt.y - kp.pt.y; } double cost = 0, lastCost = 0; bool succ = true; // indicate if this point succeeded // Gauss-Newton iterations Eigen::Matrix2d H = Eigen::Matrix2d::Zero(); // hessian Eigen::Vector2d b = Eigen::Vector2d::Zero(); // bias Eigen::Vector2d J; // jacobian for (int iter = 0; iter &lt; iterations; iter++) { if (inverse == false) { H = Eigen::Matrix2d::Zero(); b = Eigen::Vector2d::Zero(); } else { // only reset b b = Eigen::Vector2d::Zero(); } cost = 0; // compute cost and jacobian for (int x = -half_patch_size; x &lt; half_patch_size; x++) for (int y = -half_patch_size; y &lt; half_patch_size; y++) { double error = GetPixelValue(img1, kp.pt.x + x, kp.pt.y + y) - GetPixelValue(img2, kp.pt.x + x + dx, kp.pt.y + y + dy);; // Jacobian if (inverse == false) { J = -1.0 * Eigen::Vector2d( 0.5 * (GetPixelValue(img2, kp.pt.x + dx + x + 1, kp.pt.y + dy + y) - GetPixelValue(img2, kp.pt.x + dx + x - 1, kp.pt.y + dy + y)), 0.5 * (GetPixelValue(img2, kp.pt.x + dx + x, kp.pt.y + dy + y + 1) - GetPixelValue(img2, kp.pt.x + dx + x, kp.pt.y + dy + y - 1)) ); } else if (iter == 0) { // in inverse mode, J keeps same for all iterations // NOTE this J does not change when dx, dy is updated, so we can store it and only compute error J = -1.0 * Eigen::Vector2d( 0.5 * (GetPixelValue(img1, kp.pt.x + x + 1, kp.pt.y + y) - GetPixelValue(img1, kp.pt.x + x - 1, kp.pt.y + y)), 0.5 * (GetPixelValue(img1, kp.pt.x + x, kp.pt.y + y + 1) - GetPixelValue(img1, kp.pt.x + x, kp.pt.y + y - 1)) ); } // compute H, b and set cost; b += -error * J; cost += error * error; if (inverse == false || iter == 0) { // also update H H += J * J.transpose(); } } // compute update Eigen::Vector2d update = H.ldlt().solve(b); if (std::isnan(update[0])) { // sometimes occurred when we have a black or white patch and H is irreversible cout &lt;&lt; \"update is nan\" &lt;&lt; endl; succ = false; break; } if (iter &gt; 0 &amp;&amp; cost &gt; lastCost) { break; } // update dx, dy dx += update[0]; dy += update[1]; lastCost = cost; succ = true; if (update.norm() &lt; 1e-2) { // converge break; } } success[i] = succ; // set kp2 kp2[i].pt = kp.pt + Point2f(dx, dy); }}需要注意的是，LK光流技术有一个较强的假设：要求图像的亮度恒定，就是同一点随着时间的变化，其亮度不会发生改变。因此，相比于SIFT等特征提取算法而言，LK的速度较快，但损失了一定的精度和鲁棒性，在具体使用的过程中还是要见仁见智。要想对SLAM有深入的研究，处理cv相关的代码能力外，也少不了扎实的数学基础。例如，四元数、李群李代数（尤其是SO(3)和SE(3)这些特殊的群)。同时，还需要有一定的图形学基础和非线性优化的能力。在不断探索的过程中，后端优化的Bundle Adjustment与Loop Closure都亟需丰富的统计学知识和实操上手的优化能力，不管是传统的KF、EKF还是流行的NLO，限于篇幅这篇博客不能一一揽括，因此就集中先把前端技术的思考记录下来，综合成一篇文字，以飨读者。Acknowledgements and References Liu H , Zhang G , Bao H . A survey of monocular simultaneous localization and mapping[J]. Journal of Computer-Aided Design &amp; Computer Graphics, 2016. &#8617; Smith, Randall &amp; Cheeseman, Peter. (1987). On the Representation and Estimation of Spatial Uncertainty. The International Journal of Robotics Research. 5. 10.1177/027836498600500404. &#8617; G. Yang, Y. Wang, J. Zhi, W. Liu, Y. Shao and P. Peng, “A Review of Visual Odometry in SLAM Techniques,” 2020 International Conference on Artificial Intelligence and Electromechanical Automation (AIEA), 2020, pp. 332-336, doi: 10.1109/AIEA51086.2020.00075. &#8617; B. X. Hon, H. Tian, F. Wang, B. M. Chen and T. H. Lee, “A customized fastslam algorithm using scanning laser range finder in structured indoor environments,” 2013 10th IEEE International Conference on Control and Automation (ICCA), 2013, pp. 640-645, doi: 10.1109/ICCA.2013.6565202. &#8617; Trajkovic, Miroslav &amp; Hedley, Mark. (1998). Fast Corner Detection. Image and Vision Computing. 16. 75-87. 10.1016/S0262-8856(97)00056-5. &#8617; R. Mur-Artal, J. M. M. Montiel and J. D. Tardós, “ORB-SLAM: A Versatile and Accurate Monocular SLAM System,” in IEEE Transactions on Robotics, vol. 31, no. 5, pp. 1147-1163, Oct. 2015, doi: 10.1109/TRO.2015.2463671. &#8617; &#8617;2 &#8617;3 Baker S , Matthews I . Lucas-Kanade 20 Years On: A Unifying Framework[J]. International Journal of Computer Vision, 2004, 56(3):221-255. &#8617;"
    } ,
  
    {
      "title"       : "Database Review 1",
      "category"    : "",
      "tags"        : "database, coding",
      "url"         : "./DBRivision-1.html",
      "date"        : "2022-03-03 16:32:20 +0800",
      "description" : "A brief review to database.",
      "content"     : "DBRiview-1 数据模型概念模型及其作用 from wiki:The conceptual level unifies the various external views into a compatible global view.[36] It provides the synthesis of all the external views. It is out of the scope of the various database end-users, and is rather of interest to database application developers and database administrators. 概念模型实际上是现实世界到机器世界的一个中间层次。概念模型用于信息世界的建模，是现实世界到信息世界的第一层抽象，是数据库设计人员进行数据库设计的有力工具，也是数据库设计人员和用户之间进行交流的语言。概念模型常用E-R图表示，之后会给出实例实体： 客观存在并可相互区别的事物称为实体。实体可以是具体的人、事、物·也可以是抽象的概念或联系，例如，一个职工、一个学生、一个部门、一门课、学生的一次选课、部 门的一次订货、教师与院系的工作关系（即某位教师在某院系工作）等都是实体。实体型： 具有相同属性的实体必然具有共同的特征和性质。用实体名及其属性名集合来抽象和刻画同类实体，称为实体型。例如，学生（学号，姓名，性别，出生年月，所在院系，入 学时间）就是一个实体型。实体集: ​ 同一类型实体的集合称为实体集。例如，全体学生就是一个实体集。学生作为类别是实体型实体之间的联系： 在现实世界中，事物内部以及事物之间是有联系的，这些联系在信息世界中反映为实体（型）内部的联系和实体（型）之间的联系。实体内部的联系通常是指组成实体的各属 性之间的联系，实体之间的联系通常是指不同实体集之间的联系。 实体之间的联系有一对一、一对多和多对多等多种类型。概念模型实例： 学校中有若干系，每个系有若干班级和教研室，每个教研室有若干教员，其中有的教授和副教授 每人各带若干研究生，每个班有若干学生，每个学生选修若干课程，每门课可由若干学生选修。请用E-R 图画出此学校的概念模型。 某工厂生产若干产品，每种产品由不同的零件组成，有的零件可用在不同的产品上。这些零件 由不同的原材料制成，不同零件所用的材料可以相同。这些零件按所属的不同产品分别放在仓库中， 原材料按照类别放在若干仓库中。请用E-R图画出此工厂产品、零件、材料、仓库的概念模型。 现开发⼀套销售管理系统，需保存交易记录信息，包括销售⼈员身份证号、顾客身份证号、售卖货品名 称、数量、单价。请绘制数据库建模的ER图。 现开发⼀套销售管理系统，需保存进销存信息，包括：1). 货品清单，包括货品编号、货品名称、单价、 库存数量；2). 交易记录，包括销售⼈员身份证号、顾客身份证号、售卖货品编号。请绘制数据库建模的 ER图。 现开发⼀套销售管理系统，需保存进销存信息，包括：1). 货品清单，包括货品编号、货品名称、单价、 库存数量；2). 交易记录，包括销售⼈员身份证号、顾客身份证号、售卖货品编号。请绘制数据库建模的 ER图。 现开发⼀套销售管理系统，需保存进销存信息，包括：1). 货品清单，包括货品编号、货品名称、单价、 库存数量；2). 销售⼈员信息，包括⼈员身份证号，姓名，性别，职级，薪⽔； 3). 顾客信息，包括身份 证号，姓名，会员卡号，⽣⽇； 4). 交易记录，包括销售⼈员身份证号、顾客身份证号、售卖货品编号。 请绘制数据库建模的ER图。 Acknowledgments 数据库系统概论（第五版），王珊，萨师煊等 数据库系统概论学习指导与习题解答，王珊等 数据库系统概念（中文第六版） Wikipedia"
    } ,
  
    {
      "title"       : "Database Review 0",
      "category"    : "",
      "tags"        : "database, coding",
      "url"         : "./DBRivision-0.html",
      "date"        : "2022-03-03 13:32:20 +0800",
      "description" : "A brief review to database.",
      "content"     : "DBRiview-0 数据、数据库数据、数据库、数据库管理系统、数据库系统 from wiki:In computing, a database is an organized collection of data stored and accessed electronically. Small databases can be stored on a file system, while large databases are hosted on computer clusters or cloud storage. The design of databases spans formal techniques and practical considerations including data modeling, efficient data representation and storage, query languages, security and privacy of sensitive data, and distributed computing issues including supporting concurrent access and fault tolerance. A database management system (DBMS) is the software that interacts with end users, applications, and the database itself to capture and analyze the data. The DBMS software additionally encompasses the core facilities provided to administer the database. The sum total of the database, the DBMS and the associated applications can be referred to as a database system. Often the term “database” is also used loosely to refer to any of the DBMS, the database system or an application associated with the database.数据： 可以对数据做如下定义：描述事物的符号记录称为数据。描述事物的符号可以是数字，也可以是文字、图形、图像、音频、视频等，数据有多种表现形式，它们都可以经过数字化后存入计算机。数据库（Database）： 存放数据的仓库，这个仓库是在计算机存储设备上，而且数据是一定的格式存放的。 严格地讲，数据库是长期储存在计算机内、有组织的、可共享的大量数据的集合。数据库中的数据按一定的数据模型组织、描述和储存，（按照关系数据模型的为关系数据库）具有较小的冗余度(redundancy)、较高的数据独立性（data independency）和易扩展性（scalability），并可为各种用户共享。数据库管理系统（Database Management System）: ​ 一种系统软件，是数据库系统的核心。主要包含以下功能： 数据定义功能数据库管理系统提供数据定义语言（Data Definition Language，DDL)，用户通过它可以方便地对数据库中的数据对象的组成与结构进行定义。 数据组织、存储和管理 数据库管理系统要分类组织、存储和管理各种数据，包括数据字典、用户数据、数 的存取路径等。要确定以何种文件结构和存取方式在存储级上组织这些数据，如何实 覲数据之间的联系。数据组织和存储的基本目标是提高存储空间利用率和方便存取，提 哄多种存取方法（如索引查找、hash查找、顺序查找等）来提高存取效率。 数据操纵功能 数据库管理系统还提供数据操纵语言（Data Manipulation Language，DML)，用户可以 使用它操纵数据，实现对数据库的基本操作，如查询、插入、删除和修改等。 数据库的事务管理和运行管理 数据库在建立、运用和维护时由数据库管理系统统一管理和控制，以保证事务的正确 运行，保证数据的安全性、完整性、多用户对数据的并发使用及发生故障后的系统恢复。 数据库的建立和维护功能 数据库的建立和维护功能包括数据库初始数据的输入、转换功能，数据库的转储、恢 夏功能，数据库的重组织功能和性能监视、分析功能等。这些功能通常是由一些实用程序 管理工具 数据库系统： 数据库系统是由数据库、数据库管理系统（及其应用开发工具）、应用程序和数据库 管理员（Database Administrator，DBA)组成的存储、管理、处理和维护数据的系统。在一般不引起混淆的情况下，人们常常把数据库系统简称为数据库。文件系统与数据库系统 用文件系统管理数据具有如下特点： 数据可以长期保存 由于计算机大量用于数据处理，数据需要长期保留在外存上反复进行查询、修改、插 入和删除等操作。 由文件系统管理数据 由专门的软件即文件系统进行数据管理，文件系统把数据组织成相互独立的数据文 件，利用“按文件名访问（考虑inode)，按记录进行存取”的管理技术，提供了对文件进行打开与关闭、 对记录读取和写入等存取方式。文件系统实现了记录内的结构性。 但是，文件系统仍存在以下缺点： 数据共享性差，冗余度大 在文件系统中，一个（或一组）文件基本上对应于一个应用程序，即文件仍然是面向应用的。当不同的应用程序具有部分相同的数据时，也必须建立各自的文件，而不能共享 相同的数据，因此数据的冗余度大，浪费存储空间。同时由于相同数据的重复存储、各自 管理，容易造成数据的不一致性，给数据的修改和维护带来了困难。 数据独立性差 文件系统中的文件是为某一特定应用服务的，文件的逻辑结构是针对具体的应用来设计和优化的，因此要想对文件中的数据再增加一些新的应用会很困难。而且，当数据的逻辑结构改变时，应用程序中文件结构的定义必须修改，应用程序中对数据的使用也要改变， 因此数据依赖于应用程序，缺乏独立性。可见，文件系统仍然是一个不具有弹性的无整体结构的数据集合，即文件之间是孤立的，不能反映现实世界事物之间的内在联系。 相对于文件系统而言，数据库系统具有以下特点： 数据结构化 数据库系统实现整体数据的结构化，这是数据库的主要特征之一，也是数据库系统与 文件系统的本质区别。 在文件系统中，文件中的记录内部具有结构，但是记录的结构和记录之间的联系被固化在程序中，需要由程序员加以维护。这种工作模式既加重了程序员的负担，又不利于结构的变动。 所谓“整体”结构化是指数据库中的数据不再仅仅针对某一个应用，而是面向整个组织或企业：不仅数据内部是结构化的，而且整体是结构化的，数据之间是具有联系的考虑schema。 数据的共享性高、冗余度低且易扩充 数据库系统从整体角度看待和描述数据，数据不再面向某个应用而是面向整个系统， 因此数据可以被多个用户、多个应用共享使用。数据共享可以大大减少数据冗余，节约存 储空间。数据共享还能够避免数据之间的不相容性与不一致性。 所谓数据的不一致性是指同一数据不同副本的值不一样ACID。采用人工管理或文件系统管理时，由于数据被重复存储，不同的应用不一致。在数据库中数据共享减少了由于数据冗余造成的不一致现象。 由于数据面向整个系统，是有结构的数据，不仅可以被多个应用共享使用，而且容易增加新的应用，这就使得数据库系统弹性大，易于扩充，可以适应各种用户的要求。可以 选取整体数据的各种子集用于不同的应用系统，当应用需求改变或增加时，只要重新选取不同的子集或加上一部分数据便可以满足新的需求。 数据独立性高 数据独立性是借助数据库管理数据的一个显著优点，通常由DBMS的二级印象保证。它己成为数据库领域中一个常用术语和重要概念，包括数据的物理独立性和逻辑独立性。 物理独立性是指用户的应用程序与数据库中数据的物理存储是相互独立的。也就是说，数据在数据库中怎样存储是山数据库管理系统管理的，用户程序不需要了解，应用程序要处理的只是数据的逻辑结构，这样当数据的物理存储改变时应用程序不用改变。 逻辑独立性是指用户的应用程序与数据库的逻辑结构是相互独立的。也就是说，数据 的逻辑结构改变时用户程序也可以不变。 数据由数据库管理系统统一管理和控制 数据库的共享将会带来数据库的安全隐患，而数据库的共享是并发的（concurrency）共 享，（数据的一致性）即多个用户可以同时存取数据库中的数据，甚至可以同时存取数据库中同一个数 据，这又会带来不同用户间相互干扰的隐患。另外，数据库中数据的正确与一致也必须得到保障。为此，数据库管理系统还必须提供以下几方面的数据控制功能。 数据的安全性（security）保护数据的安全性是指保护数据以防止不合法使用造成的数据泄密和破坏。每个用户只能 按规定对某些数据以某些方式进行使用和处理。 grant or revoke--dcl 数据的完整性（integrity）检查 数据的完整性指数据的正确性、有效性和相容性。完整胜检查将数据控制在有效的范围内，并保证数据之间满足一定的关系。 并发（concurrency)控制 当多个用户的并发进程同时存取、修改数据库时，可能会发生相互干扰而得到错误的 结果或使得数据库的完整性遭到破坏，因此必须对多用户的并发操作加以控制和协调。 by transaction 数据库恢复（recovery） 计算机系统的硬件故障、软件故障、操作员的失误以及故意破坏也会影响数据库中数据的正确性，甚至造成数据库部分或全部数据的丢失。数据库管理系统必须具有将数据库 从错误状态恢复到某一己知的正确状态（亦称为完整状态或一致状态）的功能，这就是数据库的恢复功能。 rollback/undo/redo 常见的适用于文件系统和数据库系统的场景 文件系统： 包含大量半结构化数据与非结构化数据，如网页、图片等。 需要完全隔离权限的操作系统应用（通过权限矩阵控制访问），如FAT、log等（这个举例可能不太妥当） 数据库系统： 需要反复查询检索的web应用，如12306等 需要DCL保护的应用，如ERP 的不同部门等 包含大量并发性的访问与请求，如教务系统（ 数据库管理系统的主要功能： Codd proposed the following functions and services a fully-fledged general purpose DBMS should provide:[25] FROM Wiki Data storage, retrieval and update 存储、恢复、更新 User accessible catalog or data dictionary describing the metadata 如何处理columns？ Support for transactions and concurrency 并发控制 Facilities for recovering the database should it become damaged 灾备 Support for authorization of access and update of data 权限 Access support from remote locations 远程 Enforcing constraints to ensure data in the database abides by certain rules 数据约束 by the way, Edgar Frank “Ted” Codd (19 August 1923 – 18 April 2003) was an English computer scientist who, while working for IBM, invented the relational model for database management, the theoretical basis for relational databases and relational database management systems. He made other valuable contributions to computer science, but the relational model, a very influential general theory of data management, remains his most mentioned, analyzed and celebrated achievement.数据库的模式结构 上回书说到，The relational model was introduced by E.F. Codd in 1970[2] as a way to make database management systems more independent of any particular application. It is a mathematical model defined in terms of predicate logic and set theory, and implementations of it have been used by mainframe, midrange and microcomputer systems. 数据库系统的三级模式结构是指数据库系统是由外模式、模式和内模式三级构成。 模式（schema) 模式也称逻辑模式，是数据库中全体数据的逻辑结构和特征的描述，是所有用户的公共数据视图。它是数据库系统模式结构的中间层，既不涉及数据的物理存储细节和硬件环境，又与具体的应用程序、所使用的应用开发工具及高级程序设计语言无关。 模式实际上是数据库数据在逻辑级上的视图。一个数据库只有一个模式。数据库模式 以某一种数据模型为基础，统一综合地考虑了所有用户的需求，并将这些需求有机地结合 成一个逻辑整体。定义模式时不仅要定义数据的逻辑结构，例如数据记录由哪些数据项构 成，数据项的名字、类型、取值范围等；而且要定义数据之间的联系，定义与数据有关的安全性、完整性要求。 数据库管理系统提供模式数据定义语言（模式DDL)来严格地定义模式。 外模式(external schema) 外模式也称子模式（subschema）或用户模式，它是数据库用户（包括应用程序员和最 终用户）能够看见和使用的局部数据的逻辑结构和特征的描述，是数据库用户的数据视图， 是与某一应用有关的数据的逻辑表示。 外模式通常是模式的子集。一个数据库可以有多个外模式。 内模式(internal schema) 内模式也称存储模式（storage schema)，一个数据库只有一个内模式。它是数据物理 结构和存储方式的描述，是数据在数据库内部的组织方式。例如，记录的存储方式是堆存 储还是按照某个（些）属性值的升（降）序存储，或按照属性值聚簇（cluster)存储；索 引按照什么方式组织，是树索引还是hash索引：数据是否压缩存储，是否加密；数据 的存储记录结构有何规定，如定长结构或变长结构，一个记录不能跨物理页存储；等等。 数据独立性 上回书说到，数据独立性是借助数据库管理数据的一个显著优点，通常由DBMS的二级印象保证。它己成为数据库领域中一个常用术语和重要概念，包括数据的物理独立性和逻辑独立性。 物理独立性是指用户的应用程序与数据库中数据的物理存储是相互独立的。也就是说，数据在数据库中怎样存储是山数据库管理系统管理的，用户程序不需要了解，应用程序要处理的只是数据的逻辑结构，这样当数据的物理存储改变时应用程序不用改变。 逻辑独立性是指用户的应用程序与数据库的逻辑结构是相互独立的。也就是说，数据的逻辑结构改变时用户程序也可以不变。数据与程序之间的独立性使得数据的定义和描述可以从应用程序中分离出去。另外， 由于数据的存取由数据库管理系统管理，从而简化了应用程序的编制，大大减少了应用程 序的维护和修改。数据库系统的组成 数据库系统是由数据库、数据库管理系统（及其应用开发工具）、应用程序和数据库管理员（Database Administrator，DBA)组成的存储、管理、处理和维护数据的系统。实际上，数据库系统也包括支持系统的硬件及软件，以及数据库设计者及用户等人员组成sAcknowledgments 数据库系统概论，王珊，萨师煊等 数据库系统概念 Wikipedia"
    } ,
  
    {
      "title"       : "Who owns the copyright for an AI generated creative work?",
      "category"    : "opinion",
      "tags"        : "copyright, creativity, neural networks, machine learning, artificial intelligence",
      "url"         : "./AI-and-intellectual-property.html",
      "date"        : "2021-04-20 00:00:00 +0800",
      "description" : "As neural networks are used more and more in the creative process, text, images and even music are now created by AI, but who owns the copyright for those works?",
      "content"     : "Recently I was reading an article about a cool project that intends to have a neural network create songs of the late club of the 27 (artists that have tragically died at age 27 or near, and in the height of their respective careers), artists such as Amy Winehouse, Jimmy Hendrix, Curt Cobain and Jim Morrison.The project was created by Over the Bridge, an organization dedicated to increase awareness on mental health and substance abuse in the music industry, trying to denormalize and remove the glamour around such illnesses within the music community.They are using Google’s Magenta, which is a neural network that precisely was conceived to explore the role of machine learning within the creative process. Magenta has been used to create a brand new “Beatles” song or even there was a band that used it to write a full album in 2019.So, while reading the article, my immediate thought was: who owns the copyright of these new songs?Think about it, imagine one of this new songs becomes a massive hit with millions of youtube views and spotify streams, who can claim the royalties generated?At first it seems quite simple, Over the Bridge should be the ones reaping the benefits, since they are the ones who had the idea, gathered the data and then fed the neural network to get the “work of art”. But in a second thought, didn’t the original artists provide the basis for the work the neural network generated? shouldn’t their state get credit? what about Google whose tool was used, should they get credit too?Neural networks have been also used to create poetry, paintings and to write news articles, but how do they do it? A computer program developed for machine learning purposes is an algorithm that “learns” from data to make future decisions. When applied to art, music and literary works, machine learning algorithms are actually learning from some input data to generate a new piece of work, making independent decisions throughout the process to determine what the new work looks like. An important feature of this is that while programmers can set the parameters, the work is actually generated by the neural network itself, in a process akin to the thought processes of humans.Now, creative works qualify for copyright protection if they are original, with most definitions of originality requiring a human author. Most jurisdictions, including Spain and Germany, specifically state that only works created by a human can be protected by copyright. In the United States, for example, the Copyright Office has declared that it will “register an original work of authorship, provided that the work was created by a human being.”So as we currently stand, a human author is required to grant a copyright, which makes sense, there is no point of having a neural network be the beneficiary of royalties of a creative work (no bank would open an account for them anyways, lol).I think amendments have to be made to the law to ensure that the person who undertook all the arrangements necessary for the work to be created by the neural network gets the credit but also we need to modify copyright law to ensure the original authors of the body of work used as data input to produce the new piece get their corresponding share of credit. This will get messy if someone uses for example the #1 song of every month in a decade to create the decade song, then there would be as many as 120 different artists to credit.In a computer generated artistic work, both the person who undertook all the arrangements necessary for its creation as well as the original authors of the data input need to be credited.There will still be some ambiguity as to who undertook the arrangements necessary, only the one who gathered the data and pressed the button to let the network learn, or does the person who created the neural network’s model also get credit? Shall we go all the way and say that even the programmer of the neural network gets some credit as well?There are some countries, in particular the UK where some progress has been made to amend copyright laws to cater for computer generated works of art, but I believe this is one of those fields where technology will surpass our law making capacity and we will live under a grey area for a while, and maybe this is just what we need, by having these works ending up free for use by anyone in the world, perhaps a new model for remunerating creative work can be established, one that does not require commercial success to be necessary for artists to make a living, and thus they can become free to explore their art.Perhaps a new model for remunerating creative work can be established, one that does not require commercial success to be necessary for artists to make a living.The Next Rembrandt is a computer-generated 3-D–printed painting developed by a facial-recognition algorithm that scanned data from 346 known paintings by the Dutch painter in a process lasting 18 months. The portrait is based on 168,263 fragments from Rembrandt’s works."
    } ,
  
    {
      "title"       : "So, what is a neural network?",
      "category"    : "theory",
      "tags"        : "neural networks, machine learning, artificial intelligence",
      "url"         : "./back-to-basics.html",
      "date"        : "2021-04-02 00:00:00 +0800",
      "description" : "ELI5: what is a neural network.",
      "content"     : "The omnipresence of technology nowadays has made it commonplace to read news about AI, just a quick glance at today’s headlines, and I get: This Powerful AI Technique Led to Clashes at Google and Fierce Debate in Tech. How A.I.-powered companies dodged the worst damage from COVID AI technology detects ‘ticking time bomb’ arteries AI in Drug Discovery Starts to Live Up to the Hype Pentagon seeks commercial solutions to get its data ready for AITopics from business, manufacturing, supply chain, medicine and biotech and even defense are covered in those news headlines, definitively the advancements on the fields of artificial intelligence, in particular machine learning and deep neural networks have permeated into our daily lives and are here to stay. But, do the general population know what are we talking about when we say “an AI”? I assume most people correctly imagine a computer algorithm or perhaps the more adventurous minds think of a physical machine, an advanced computer entity or even a robot, getting smarter by itself with every use-case we throw at it. And most people will be right, when “an AI” is mentioned it is indeed an algorithm run by a computer, and there is where the boundary of their knowledge lies.They say that the best way to learn something is to try to explain it, so in a personal exercise I will try to do an ELI5 (Explain it Like I am 5) version of what is a neural network.Let’s start with a little history, humans have been tinkering with the idea of an intelligent machine for a while now, some even say that the idea of artificial intelligence was conceived by the ancient greeks (source), and several attempts at devising “intelligent” machines have been made through history, a notable one was ‘The Analytical Engine’ created by Charles Babbage in 1837:The Analytical Engine of Charles Babbage - 1837Then, in the middle of last century by trying to create a model of how our brain works, Neural Networks were born. Around that time, Frank Rosenblatt at Cornell trying to understand the simple decision system present in the eye of a common housefly, proposed the idea of a perceptron, a very simple system that processes certain inputs with basic math operations and produces an output.To illustrate, let’s say that the brain of the housefly is a perceptron, its inputs are whatever values are produced by the multiple cells in its eyes, when the eye cell detects “something” it’s output will be a 1, and if there is nothing a 0. Then the combination of all those inputs can be processed by the perceptron (the fly brain), and the output is a simple 0 or 1 value. If it is a 1 then the brain is telling the fly to flee and if it is a 0 it means it is safe to stay where it is.We can imagine then that if many of the eye cells of the fly produce 1s, it means that an object is quite near, and therefore the perceptron will calculate a 1, it is time to flee.The perceptron is just a math operation, one that multiplies certain input values with preset “parameters” (called weights) and adds up the resulting multiplications to generate a value.Then the magic spark was ignited, the parameters (weights) of the perceptron could be “learnt” by a process of minimizing the difference between known results of particular observations, and what the perceptron is actually calculating. It is this process of learning what we call training the neural network.This idea is so powerful that even today it is one of the fundamental building blocks of what we call AI.From this I will try to explain how this simple concept can have such diverse applications as natural language processing (think Alexa), image recognition like medical diagnosis from a CTR scan, autonomous vehicles, etc.A basic neural network is a combination of perceptrons in different arrangements, the perceptron therefore was downgraded from “fly brain” to “network neuron”.A neural network has different components, in its basic form it has: Input Hidden layers OutputInputThe inputs of a neural network are in their essence just numbers, therefore anything that can be converted to a number can become an input. Letters in a text, pixels in an image, frequencies in a sound wave, values from a sensor, etc. are all different things that when converted to a numerical value serve as inputs for the neural network. This is one of the reasons why applications of neural networks are so diverse.Inputs can be as many as one need for the task at hand, from maybe 9 inputs to teach a neural network how to play tic-tac-toe to thousands of pixels from a camera for an autonomous vehicle. Since the input of a perceptron needs to be a single value, if for example a color pixel is chosen as input, it most likely will be broken into three different values; its red, green and blue components, hence each pixel will become 3 different inputs for the neural network.Hidden layersA “layer” within a neural network is just a group of perceptrons that all perform the same exact mathematical operation to the inputs and produce an output. The catch is that each of them have different weights (parameters), therefore their output for a given input will be different amongst them. There are many types of layers, the most typical of them being a “dense” layer, which is another word to say that all the inputs are connected to all the neurons (individual perceptrons), and as said before, each of these connections have a weight associated with it, so that the operation that each neuron performs is a simple weighted sum of all the inputs.The hidden layer is then typically connected to another dense layer, and their connection means that each output of a neuron from the first layer is treated effectively as an input for the subsequent one, and it is thus connected to every neuron.A neural network can have from one to as many layers as one can think, and the number of layers depends solely on the experience we have gathered on the particular problem we would like to solve.Another critical parameter of a hidden layer is the number of neurons it has, and again, we need to rely on experience to determine how many neurons are needed for a given problem. I have seen networks that vary from a couple of neurons to the thousands. And of course each hidden layer can have as many neurons as we please, so the number of combinations is vast.To the number of layers, their type and how many neurons each have, is what we call the network topology (including the number of inputs and outputs).OutputAt the very end of the chain, another layer lies (which behaves just like a hidden layer), but has the peculiarity that it is the final layer, and therefore whatever it calculates will be the output values of the whole network. The number of outputs the network has is a function of the problem we would like to solve. It could be as simple as one output, with its value representing a probability of an action (like in the case of the flee reaction of the housefly), to many outputs, perhaps if our network is trying to distinguish images of animals, one would have an output for each animal species, and the output would represent how much confidence the network has that the particular image belongs to the corresponding species.As we said, the neural network is just a collection of individual neurons, doing basic math operations on certain inputs in series of layers that eventually generate an output. This mesh of neurons is then “trained” on certain output values from known cases of the inputs; once it has learned it can then process new inputs, values that it has never seen before with surprisingly accurate results.Many of the problems neural networks solve, could be certainly worked out by other algorithms, however, since neural networks are in their core very basic operations, once trained, they are extremely efficient, hence much quicker and economical to produce results.There are a few more details on how a simple neural network operate that I purposedly left out to make this explanation as simple as possible. Thinks like biases, the activation functions and the math behind learning, the backpropagation algorithm, I will leave to a more in depth article. I will also write (perhaps in a series) about the more complex topologies combining different types of layers and other building blocks, a part from the perceptron.Things like “Alexa”, are a bit more complex, but work on exactly the same principles. Let’s break down for example the case of asking “Alexa” to play a song in spotify. Alexa uses several different neural networks to acomplish this:1. Speech recognitionAs a basic input we have our speech: the command “Alexa, play Van Halen”. This might seem quite simple for us humans to process, but for a machine is an incredible difficult feat to be able to understand speech, things like each individual voice timbre, entonation, intention and many more nuances of human spoken language make it so that traditional algorithms have struggled a lot with this. In our simplified example let’s say that we use a neural network to transform our spoken speech into text characters a computer is much more familiarized to learn.2. Understanding what we mean (Natural Language Understanding)Once the previous network managed to succesfuly convert our spoken words into text, there comes the even more difficult task of making sense of what we said. Things that we humans take for granted such as context, intonation and non verbal communication, help give our words meaning in a very subtle, but powerful way, a machine will have to do with much less information to correctly understand what we mean. It has to correctly identify the intention of our sentence and the subject or entities of what we mean.The neural network has to identify that it received a command (by identifying its name), the command (“play music”), and our choice (“Van Halen”). And it does so by means of simple math operations as described before. Of course the network involved is quite complex and has different types of neurons and connection types, but the underlying principles remain.3. Replying to usOnce Alexa understood what we meant, it then proceeds to execute the action of the command it interpreted and it replies to us in turn using natural language. This is accomplished using a technique called speech synthesis, things like pitch, duration and intensity of the words and phonems are selected based on the “meaning” of what Alexa will respond to us: “Playing songs by Van Halen on Spotify” sounding quite naturally. And all is accomplished with neural networks executing many simple math operations.Although it seems quite complex, the process for AI to understand us can be boiled down to simple math operationsOf course Amazon’s Alexa neural networks have undergone quite a lot of training to get to the level where they are, the beauty is that once trained, to perform their magic they just need a few mathematical operations.As said before, I will continue to write about the basics of neural networks, the next article in the series will dive a bit deeper into the math behind a basic neural network."
    } ,
  
    {
      "title"       : "Starting the adventure",
      "category"    : "",
      "tags"        : "general blogging, thoughts, life",
      "url"         : "./starting-the-adventure.html",
      "date"        : "2021-03-24 00:00:00 +0800",
      "description" : "Midlife career change: a disaster or an opportunity?",
      "content"     : "In the midst of a global pandemic caused by the SARS-COV2 coronavirus; I decided to start blogging. I wanted to blog since a long time, I have always enjoyed writing, but many unknowns and having “no time” for it prevented me from taking it up. Things like: “I don’t really know who my target audience is”, “what would my topic or topics be?”, “I don’t think I am a world-class expert in anything”, and many more kept stopping me from setting up my own blog. Now seemed like a good time as any so with those and tons of other questions in my mind I decided it was time to start.Funnily, this is not my first post. The birth of the blog came very natural as a way to “document” my newly established pursuit for getting myself into Machine Learning. This new adventure of mine comprises several things, and if I want to succeed I need to be serious about them all: I want to start coding again! I used to code a long time ago, starting when I was 8 years old in a Tandy Color Computer hooked up to my parent’s TV. Machine Learning is a vast, wide subject, I want to learn the generals, but also to select a few areas to focus on. Setting up a blog to document my journey and share it: Establish a learning and blogging routine. If I don’t do this, I am sure this endeavour will die off soon.As for the focus areas I will start with: Neural Networks fundamentals: history, basic architecture and math behind them Deep Neural Networks Reinforcement Learning Current state of the art: what is at the cutting edge now in terms of Deep Neural Networks and Reinforcement Learning?I selected the above areas to focus on based on my personal interests, I have been fascinated by the developments in reinforcement learning for a long time, in particular Deep Mind’s awesome Go, Chess and Starcraft playing agents. Therefore, I started reading a lot about it and even started a personal project for coding a tic-tac-toe learning agent.With my limited knowledge I have drafted the following learning path: Youtube: Three Blue One Brown’s videos on Neural Networks, Calculus and Linear Algebra. I cannot recommend them enough, they are of sufficient depth and use animation superbly to facilitate the understanding of the subjects. Coursera: Andrew Ng’s Machine Learning course Book: Deep Learning with Python by Francois Chollet Book: Reinforcement Learning: An Introduction, by Richard S. Sutton and Andrew G. BartoAs for practical work I decided to start by coding my first models from scratch (without using libraries such as Tensorflow), to be able to deeply understand the math and logic behind the models, so far it has proven to be priceless.For my next project I think I will start to do the basic hand-written digits recognition, which is the Machine Learning Hello World, for this I think I will start to use Tensorflow already.I will continue to write about my learning road, what I find interesting and relevant, and to document all my practical exercises, as well as news and the state of the art in the world of AI.So far, all I have learned has been so engaging that I am seriously thinking of a career change. I have 17 years of international experience in multinational corporations across various functions, such as Information Services, Sales, Customer Care and New Products Introduction, and sincerely, I am finding more joy in artificial intelligence than anything else I have worked on before. Let’s see where the winds take us.Thanks for reading!P.S. For the geeks like me, here is a snippet on the technical side of the blog.Static Website GeneratorI researched a lot on this, when I started I didn’t even know I needed a static website generator. I was just sure of one thing, I wanted my blog site to look modern, be easy to update and not to have anything extra or additional content or functionality I did not need.There is a myriad of website generators nowadays, after a lengthy search the ones I ended up considering are: wordpress wix squarespace ghost webflow netlify hugo gatsby jekyllI started with the web interfaced generators with included hosting in their offerings:wordpress is the old standard, it is the one CMS I knew from before, and I thought I needed a fully fledged CMS, so I blindly ran towards it. Turns out, it has grown a lot since I remembered, it is now a fully fledged platform for complex websites and ecommerce development, even so I decided to give it a try, I picked a template and created a site. Even with the most simplistic and basic template I could find, there is a lot going on in the site. Setting it up was not as difficult or cumbersome as others claim, it took me about one hour to have it up and running, it looks good, but a bit crowded for my personal taste, and I found out it serves ads in your site for the readers, that is a big no for me.I have tried wix and squarespace before, they are fantastic for quick and easy website generation, but their free offering has ads, so again, a big no for me.I discovered ghost as the platform used by one of the bloggers I follow (Sebastian Ruder), turns out is a fantastic evolution over wordpress. It runs on the latest technologies, its interface is quite modern, and it is focused on one thing only: publishing. They have a paid hosting service, but the software is open sourced, therefore free to use in any hosting.I also tested webflow and even created a mockup there, the learning curve was quite smooth, and its CMS seems quite robust, but a bit too much for the functionalities I required.Next were the generators that don’t have a web interface, but can be easily set up:The first I tried was netlify, I also set up a test site in it. Netlify provides free hosting, and to keep your source files it uses GitHub (a repository keeps the source files where it publishes from). It has its own CMS, Netlify CMS, and you have a choice of site generators: Hugo, Gatsby, MiddleMan, Preact CLI, Next.js, Elevently and Nuxt.js, and once you choose there are some templates for each. I did not find the variety of templates enticing enough, and the set up process was much more cumbersome than with wordpress (at least for my knowledge level). I choose Hugo for my test site.I also tested gatsby with it’s own Gatsby Cloud hosting service, here is my test site. They also use GitHub as a base to host the source files to build the website, so you create a repository, and it is connected to it. I found the free template offerings quite limited for what I was looking for.Finally it came the turn for jekyll, although an older, and slower generator (compared to Hugo and Gatsby), it was created by one of the founders of GitHub, so it’s integration with GitHub Pages is quite natural and painless, so much so, that to use them together you don’t even have to install Jekyll in your machine! You have two choices: keep it all online, by having one repository in Github keep all the source files, modify or add them online, and having Jekyll build and publish your site to the special gh-pages repository everytime you change or add a new file to the source repository. Have a synchronized local copy of the source files for the website, this way you can edit your blog and customize it in your choice of IDE (Integrated Development Environment). Then, when you update any file on your computer, you just “push” the changes to GitHub, and GitHub Pages automatically uses Jekyll to build and publish your site.I chose the second option, specially because I can manipulate files, like images, in my laptop, and everytime I sync my local repository with GitHub, they are updated and published automatically. Quite convenient.After testing with several templates to get the feel for it, I decided to keep Jekyll for my blog for several reasons: the convenience of not having to install anything extra on my computer to build my blog, the integration with GitHub Pages, the ease of use, the future proofing via integration with modern technologies such as react or vue and the vast online community that has produced tons of templates and useful information for issue resolution, customization and added functionality.I picked up a template, just forked the repository and started modifying the files to customize it, it was fast and easy, I even took it upon myself to add some functionality to the template (it served as a coding little project) like: SEO meta tags Dark mode (configurable in _config.yml file) automatic sitemap.xml automatic archive page with infinite scrolling capability new page of posts filtered by a single tag (without needing autopages from paginator V2), also with infinite scrolling click to tweet functionality (just add a &lt;tweet&gt; &lt;/tweet&gt; tag in your markdown. custom and responsive 404 page responsive and automatic Table of Contents (optional per post) read time per post automatically calculated responsive post tags and social share icons (sticky or inline) included linkedin, reddit and bandcamp icons copy link to clipboard sharing option (and icon) view on github link button (optional per post) MathJax support (optional per post) tag cloud in the home page ‘back to top’ button comments ‘courtain’ to mask the disqus interface until the user clicks on it (configurable in _config.yml) CSS variables to make it easy to customize all colors and fonts added several pygments themes for code syntax highlight configurable from the _config.yml file. See the highlighter directory for reference on the options. responsive footer menu and footer logo (if setup in the config file) smoother menu animationsAs a summary, Hugo and Gatsby might be much faster than Jekyll to build the sites, but their complexity I think makes them useful for a big site with plenty of posts. For a small site like mine, Jekyll provides sufficient functionality and power without the hassle.You can use the modified template yourself by forking my repository. Let me know in the comments or feel free to contact me if you are interested in a detailed walkthrough on how to set it all up.HostingSince I decided on Jekyll to generate my site, the choice for hosting was quite obvious, Github Pages is very nicely integrated with it, it is free, and it has no ads! Plus the domain name isn’t too terrible (the-mvm.github.io).Interplanetary File SystemTo contribute to and test IPFS I also set up a mirror in IPFS by using fleek.co. I must confess that it was more troublesome than I imagined, it was definetively not plug and play because of the paths used to fetch resources. The nature of IPFS makes short absolute paths for website resources (like images, css and javascript files) inoperative; the easiest fix for this is to use relative paths, however the same relative path that works for the root directory (i.e. /index.html) does not work for links inside directories (i.e. /tags/), and since the site is static, while generating it, one must make the distinction between the different directory levels for the page to be rendered correctly.At first I tried a simple (but brute force solution):# determine the level of the current file{% assign lvl = page.url | append:'X' | split:'/' | size %}# create the relative base (i.e. \"../\"){% capture relativebase %}{% for i in (3..lvl) %}../{% endfor %}{% endcapture %}{% if relativebase == '' %} {% assign relativebase = './' %}{% endif %}...# Eliminate unecesary double backslashes{% capture post_url %}{{ relativebase }}{{ post.url }}{% endcapture %}{% assign post_url = post_url | replace: \"//\", \"/\" %}This jekyll/liquid code was executed in every page (or include) that needed to reference a resource hosted in the same server.But this fix did not work for the search function, because it relies on a search.json file (also generated programmatically to be served as a static file), therefore when generating this file one either use the relative path for the root directory or for a nested directory, thus the search results will only link correctly the corresponding pages if the page where the user searched for something is in the corresponding scope.So the final solution was to make the whole site flat, meaning to live in a single directory. All pages and posts will live under the root directory, and by doing so, I can control how to address the relative paths for resources."
    } ,
  
    {
      "title"       : "Deep Q Learning for Tic Tac Toe",
      "category"    : "",
      "tags"        : "machine learning, artificial intelligence, reinforcement learning, coding, python",
      "url"         : "./deep-q-learning-tic-tac-toe.html",
      "date"        : "2021-03-19 05:14:20 +0800",
      "description" : "Inspired by Deep Mind's astonishing feats of having their Alpha Go, Alpha Zero and Alpha Star programs learn (and be amazing at it) Go, Chess, Atari games and lately Starcraft; I set myself to the task of programming a neural network that will learn by itself how to play the ancient game of tic tac toe. How hard could it be?",
      "content"     : "BackgroundAfter many years of a corporate career (17) diverging from computer science, I have now decided to learn Machine Learning and in the process return to coding (something I have always loved!).To fully grasp the essence of ML I decided to start by coding a ML library myself, so I can fully understand the inner workings, linear algebra and calculus involved in Stochastic Gradient Descent. And on top learn Python (I used to code in C++ 20 years ago).I built a general purpose basic ML library that creates a Neural Network (only DENSE layers), saves and loads the weights into a file, does forward propagation and training (optimization of weights and biases) using SGD. I tested the ML library with the XOR problem to make sure it worked fine. You can read the blog post for it here.For the next challenge I am interested in reinforcement learning greatly inspired by Deep Mind’s astonishing feats of having their Alpha Go, Alpha Zero and Alpha Star programs learn (and be amazing at it) Go, Chess, Atari games and lately Starcraft; I set myself to the task of programming a neural network that will learn by itself how to play the ancient game of tic tac toe (or noughts and crosses).How hard could it be?Of course the first thing to do was to program the game itself, so I chose Python because I am learning it, so it gives me a good practice opportunity, and PyGame for the interface.Coding the game was quite straightforward, albeit for the hiccups of being my first PyGame and almost my first Python program ever.I created the game quite openly, in such a way that it can be played by two humans, by a human vs. an algorithmic AI, and a human vs. the neural network. And of course the neural network against a choice of 3 AI engines: random, minimax or hardcoded (an exercise I wanted to do since a long time).While training, the visuals of the game can be disabled to make training much faster.Now, for the fun part, training the network, I followed Deep Mind’s own DQN recommendations:The network will be an approximation for the Q value function or Bellman equation, meaning that the network will be trained to predict the \"value\" of each move available in a given game state.A replay experience memory was implemented. This meant that the neural network will not be trained after each move. Each move will be recorded in a special \"memory\" alongside with the state of the board and the reward it received for taking such an action (move).After the memory is sizable enough, batches of random experiences sampled from the replay memory are used for every training roundA secondary neural network (identical to the main one) is used to calculate part of the Q value function (Bellman equation), in particular the future Q values. And then it is updated with the main network's weights every n games. This is done so that we are not chasing a moving target.Designing the neural networkThe Neural Network chosen takes 9 inputs (the current state of the game) and outputs 9 Q values for each of the 9 squares in the board of the game (possible actions). Obviously some squares are illegal moves, hence while training there was a negative reward given to illegal moves hoping that the model would learn not to play illegal moves in a given position.I started out with two hidden layers of 36 neurons each, all fully connected and activated via ReLu. The output layer was initially activated using sigmoid to ensure that we get a nice value between 0 and 1 that represents the QValue of a given state action pair.The many models…Model 1 - the first tryAt first the model was trained by playing vs. a “perfect” AI, meaning a hard coded algorithm that never looses and that will win if it is given the chance. After several thousand training rounds, I noticed that the Neural Network was not learning much; so I switched to training vs. a completely random player, so that it will also learn how to win. After training vs. the random player, the Neural Network seems to have made progress and is steadily diminishing the loss function over time.However, the model was still generating many illegal moves, so I decided to modify the reinforcement learning algorithm to punish more the illegal moves. The change consisted in populating with zeros all the corresponding illegal moves for a given position at the target values to train the network. This seemed to work very well for diminishing the illegal moves:Nevertheless, the model was still performing quite poorly winning only around 50% of games vs. a completely random player (I expected it to win above 90% of the time). This was after only training 100,000 games, so I decided to keep training and see the results:Wins: 65.46% Losses: 30.32% Ties: 4.23%Note that when training restarts, the loss and illegal moves are still high in the beginning of the training round, and this is caused by the epsilon greedy strategy that prefers exploration (a completely random move) over exploitation, this preference diminishes over time.After another round of 100,000 games, I can see that the loss function actually started to diminish, and the win rate ended up at 65%, so with little hope I decided to carry on and do another round of 100,000 games (about 2 hours in an i7 MacBook Pro):Wins: 46.40% Losses: 41.33% Ties: 12.27%As you can see in the chart, the calculated loss not even plateaued, but it seemed to increase a bit over time, which tells me the model is not learning anymore. This was confirmed by the win rate decreasing with respect of the previous round to a meek 46.4% that looks no better than a random player.Model 2 - Linear activation for the outputAfter not getting the results I wanted, I decided to change the output activation function to linear, since the output is supposed to be a Q value, and not a probability of an action.Wins: 47.60% Losses: 39% Ties: 13.4%Initially I tested with only 1000 games to see if the new activation function was working, the loss function appears to be decreasing, however it reached a plateau around a value of 1, hence still not learning as expected. I came across a technique by Brad Kenstler, Carl Thome and Jeremy Jordan called Cyclical Learning Rate, which appears to solve some cases of stagnating loss functions in this type of networks. So I gave it a go using their Triangle 1 model.With the cycling learning rate in place, still no luck after a quick 1,000 games training round; so I decided to implement on top a decaying learning rate as per the following formula:The resulting learning rate combining the cycles and decay per epoch is:Learning Rate = 0.1, Decay = 0.0001, Cycle = 2048 epochs, max Learning Rate factor = 10xtrue_epoch = epoch - c.BATCH_SIZElearning_rate = self.learning_rate*(1/(1+c.DECAY_RATE*true_epoch))if c.CLR_ON: learning_rate = self.cyclic_learning_rate(learning_rate,true_epoch)@staticmethoddef cyclic_learning_rate(learning_rate, epoch): max_lr = learning_rate*c.MAX_LR_FACTOR cycle = np.floor(1+(epoch/(2*c.LR_STEP_SIZE))) x = np.abs((epoch/c.LR_STEP_SIZE)-(2*cycle)+1) return learning_rate+(max_lr-learning_rate)*np.maximum(0,(1-x))c.DECAY_RATE = learning rate decay ratec.MAX_LR_FACTOR = multiplier that determines the max learning ratec.LR_STEP_SIZE = the number of epochs each cycle lastsWith these many changes, I decided to restart with a fresh set of random weights and biases and try training more (much more) games.1,000,000 episodes, 7.5 million epochs with batches of 64 moves eachWins: 52.66% Losses: 36.02% Ties: 11.32%After 24 hours!, my computer was able to run 1,000,000 episodes (games played), which represented 7.5 million training epochs of batches of 64 plays (480 million plays learned), the learning rate did decreased (a bit), but is clearly still in a plateau; interestingly, the lower boundary of the loss function plot seems to continue to decrease as the upper bound and the moving average remains constant. This led me to believe that I might have hit a local minimum.Model 3 - new network topologyAfter all the failures I figured I had to rethink the topology of the network and play around with combinations of different networks and learning rates.100,000 episodes, 635,000 epochs with batches of 64 moves eachWins: 76.83% Losses: 17.35% Ties: 5.82%I increased to 200 neurons each hidden layer. In spite of this great improvement the loss function was still in a plateau at around 0.1 (Mean Squared Error). Which, although it is greatly reduced from what we had, still was giving out only 77% win rate vs. a random player, the network was playing tic tac toe as a toddler!*I can still beat the network most of the time! (I am playing with the red X)*100,000 more episodes, 620,000 epochs with batches of 64 moves eachWins: 82.25% Losses: 13.28% Ties: 4.46%Finally we crossed the 80% mark! This is quite an achievement, it seems that the change in network topology is working, although it also looks like the loss function is stagnating at around 0.15.After more training rounds and some experimenting with the learning rate and other parameters, I couldn’t improve past the 82.25% win rate.These have been the results so far:It is quite interesting to learn how the many parameters (hyper-parameters as most authors call them) of a neural network model affect its training performance, I have played with: the learning rate the network topology and activation functions the cycling and decaying learning rate parameters the batch size the target update cycle (when the target network is updated with the weights from the policy network) the rewards policy the epsilon greedy strategy whether to train vs. a random player or an “intelligent” AI.And so far the most effective change has been the network topology, but being so close but not quite there yet to my goal of 90% win rate vs. a random player, I will still try to optimize further.Network topology seems to have the biggest impact on a neural network's learning ability.Model 4 - implementing momentumI reached out to the reddit community and a kind soul pointed out that maybe what I need is to apply momentum to the optimization algorithm. So I did some research and ended up deciding to implement various optimization methods to experiment with: Stochastic Gradient Descent with Momentum RMSProp: Root Mean Square Plain Momentum NAG: Nezterov’s Accelerated Momentum Adam: Adaptive Moment Estimation and keep my old vanilla Gradient Descent (vGD) ☺Click here for a detailed explanation and code of all the implemented optimization algorithms.So far, I have not been able to get better results with Model 4, I have tried all the momentum optimization algorithms with little to no success.Model 5 - implementing one-hot encoding and changing topology (again)I came across an interesting project in Github that deals exactly with Deep Q Learning, and I noticed that he used “one-hot” encoding for the input as opposed to directly entering the values of the player into the 9 input slots. So I decided to give it a try and at the same time change my topology to match his:So, ‘one hot’ encoding is basically changing the input of a single square in the tic tac toe board to three numbers, so that each state is represented with different inputs, thus the network can clearly differentiate the three of them. As the original author puts it, the way I was encoding, having 0 for empty, 1 for X and 2 for O, the network couldn’t easily tell that, for instance, O and X both meant occupied states, because one is two times as far from 0 as the other. With the new encoding, the empty state will be 3 inputs: (1,0,0), the X will be (0,1,0) and the O (0,0,1) as in the diagram.Still, no luck even with Model 5, so I am starting to think that there could be a bug in my code.To test this hypothesis, I decided to implement the same model using Tensorflow / Keras.Model 6 - Tensorflow / Kerasself.PolicyNetwork = Sequential()for layer in hidden_layers: self.PolicyNetwork.add(Dense( units=layer, activation='relu', input_dim=inputs, kernel_initializer='random_uniform', bias_initializer='zeros'))self.PolicyNetwork.add(Dense( outputs, kernel_initializer='random_uniform', bias_initializer='zeros'))opt = Adam(learning_rate=c.LEARNING_RATE, beta_1=c.GAMMA_OPT, beta_2=c.BETA, epsilon=c.EPSILON, amsgrad=False)self.PolicyNetwork.compile(optimizer='adam', loss='mean_squared_error', metrics=['accuracy'])As you can see I am reusing all of my old code, and just replacing my Neural Net library with Tensorflow/Keras, keeping even my hyper-parameter constants.The training function changed to:reduce_lr_on_plateau = ReduceLROnPlateau(monitor='loss', factor=0.1, patience=25)history = self.PolicyNetwork.fit(np.asarray(states_to_train), np.asarray(targets_to_train), epochs=c.EPOCHS, batch_size=c.BATCH_SIZE, verbose=1, callbacks=[reduce_lr_on_plateau], shuffle=True)With Tensorflow implemented, the first thing I noticed, was that I had an error in the calculation of the loss, although this only affected reporting and didn’t change a thing on the training of the network, so the results kept being the same, the loss function was still stagnating! My code was not the issue.Model 7 - changing the training scheduleNext I tried to change the way the network was training as per u/elBarto015 advised me on reddit.The way I was training initially was: Games begin being simulated and the outcome recorded in the replay memory Once a sufficient ammount of experiences are recorded (at least equal to the batch size) the Network will train with a random sample of experiences from the replay memory. The ammount of experiences to sample is the batch size. The games continue to be played between the random player and the network. Every move from either player generates a new training round, again with a random sample from the replay memory. This continues until the number of games set up conclude.The first change was to train only after every game concludes with the same ammount of data (a batch). This was still not giving any good results.The second change was more drastic, it introduced the concept of epochs for every training round, it basically sampled the replay memory for epochs * batch size experiences, for instance if epochs selected were 10, and batch size was 81, then 810 experiences were sampled out of the replay memory. With this sample the network was then trained for 10 epochs randomly using the batch size.This meant that I was training now effectively 10 (or the number of epochs selected) times more per game, but in batches of the same size and randomly shuffling the experiences each epoch.After still playing around with some hyperparameters I managed to get similar performance as I got before, reaching 83.15% win rate vs. the random player, so I decided to keep training in rounds of 2,000 games each to evaluate performance. With almost every round I could see improvement:As of today, my best result so far is 87.5%, I will leave it rest for a while and keep investigating to find a reason for not being able to reach at least 90%. I read about self play, and it looks like a viable option to test and a fun coding challenge. However, before embarking in yet another big change I want to ensure I have been thorough with the model and have tested every option correctly.I feel the end is near… should I continue to update this post as new events unfold or shall I make it a multi post thread?"
    } ,
  
    {
      "title"       : "Neural Network Optimization Methods and Algorithms",
      "category"    : "",
      "tags"        : "coding, machine learning, optimization, deep Neural networks",
      "url"         : "./neural-network-optimization-methods.html",
      "date"        : "2021-03-13 03:32:20 +0800",
      "description" : "Some neural network optimization algorithms mostly to implement momentum when doing back propagation.",
      "content"     : "For the seemingly small project I undertook of creating a machine learning neural network that could learn by itself to play tic-tac-toe, I bumped into the necesity of implementing at least one momentum algorithm for the optimization of the network during backpropagation.And since my original post for the TicTacToe project is quite large already, I decided to post separately these optimization methods and how did I implement them in my code.AdamsourceAdaptive Moment Estimation (Adam) is an optimization method that computes adaptive learning rates for each weight and bias. In addition to storing an exponentially decaying average of past squared gradients \\(v_t\\) and an exponentially decaying average of past gradients \\(m_t\\), similar to momentum. Whereas momentum can be seen as a ball running down a slope, Adam behaves like a heavy ball with friction, which thus prefers flat minima in the error surface. We compute the decaying averages of past and past squared gradients \\(m_t\\) and \\(v_t\\) respectively as follows:\\(\\begin{align}\\begin{split}m_t &amp;= \\beta_1 m_{t-1} + (1 - \\beta_1) g_t \\\\v_t &amp;= \\beta_2 v_{t-1} + (1 - \\beta_2) g_t^2\\end{split}\\end{align}\\)\\(m_t\\) and \\(v_t\\) are estimates of the first moment (the mean) and the second moment (the uncentered variance) of the gradients respectively, hence the name of the method. As \\(m_t\\) and \\(v_t\\) are initialized as vectors of 0's, the authors of Adam observe that they are biased towards zero, especially during the initial time steps, and especially when the decay rates are small (i.e. \\(\\beta_1\\) and \\(\\beta_2\\) are close to 1).They counteract these biases by computing bias-corrected first and second moment estimates:\\(\\begin{align}\\begin{split}\\hat{m}_t &amp;= \\dfrac{m_t}{1 - \\beta^t_1} \\\\\\hat{v}_t &amp;= \\dfrac{v_t}{1 - \\beta^t_2} \\end{split}\\end{align}\\)We then use these to update the weights and biases which yields the Adam update rule:\\(\\theta_{t+1} = \\theta_{t} - \\dfrac{\\eta}{\\sqrt{\\hat{v}_t} + \\epsilon} \\hat{m}_t\\).The authors propose defaults of 0.9 for \\(\\beta_1\\), 0.999 for \\(\\beta_2\\), and \\(10^{-8}\\) for \\(\\epsilon\\).view on github# decaying averages of past gradientsself.v[\"dW\" + str(i)] = ((c.BETA1 * self.v[\"dW\" + str(i)]) + ((1 - c.BETA1) * np.array(self.gradients[i]) ))self.v[\"db\" + str(i)] = ((c.BETA1 * self.v[\"db\" + str(i)]) + ((1 - c.BETA1) * np.array(self.bias_gradients[i]) ))# decaying averages of past squared gradientsself.s[\"dW\" + str(i)] = ((c.BETA2 * self.s[\"dW\"+str(i)]) + ((1 - c.BETA2) * (np.square(np.array(self.gradients[i]))) ))self.s[\"db\" + str(i)] = ((c.BETA2 * self.s[\"db\" + str(i)]) + ((1 - c.BETA2) * (np.square(np.array( self.bias_gradients[i]))) ))if c.ADAM_BIAS_Correction: # bias-corrected first and second moment estimates self.v[\"dW\" + str(i)] = self.v[\"dW\" + str(i)] / (1 - (c.BETA1 ** true_epoch)) self.v[\"db\" + str(i)] = self.v[\"db\" + str(i)] / (1 - (c.BETA1 ** true_epoch)) self.s[\"dW\" + str(i)] = self.s[\"dW\" + str(i)] / (1 - (c.BETA2 ** true_epoch)) self.s[\"db\" + str(i)] = self.s[\"db\" + str(i)] / (1 - (c.BETA2 ** true_epoch))# apply to weights and biasesweight_col -= ((eta * (self.v[\"dW\" + str(i)] / (np.sqrt(self.s[\"dW\" + str(i)]) + c.EPSILON))))self.bias[i] -= ((eta * (self.v[\"db\" + str(i)] / (np.sqrt(self.s[\"db\" + str(i)]) + c.EPSILON))))SGD MomentumsourceVanilla SGD has trouble navigating ravines, i.e. areas where the surface curves much more steeply in one dimension than in another, which are common around local optima. In these scenarios, SGD oscillates across the slopes of the ravine while only making hesitant progress along the bottom towards the local optimum.Momentum is a method that helps accelerate SGD in the relevant direction and dampens oscillations. It does this by adding a fraction \\(\\gamma\\) of the update vector of the past time step to the current update vector:\\(\\begin{align}\\begin{split}v_t &amp;= \\beta_1 v_{t-1} + \\eta \\nabla_\\theta J( \\theta) \\\\\\theta &amp;= \\theta - v_t\\end{split}\\end{align}\\)The momentum term \\(\\beta_1\\) is usually set to 0.9 or a similar value.Essentially, when using momentum, we push a ball down a hill. The ball accumulates momentum as it rolls downhill, becoming faster and faster on the way (until it reaches its terminal velocity if there is air resistance, i.e. \\(\\beta_1 &lt; 1\\)). The same thing happens to our weight and biases updates: The momentum term increases for dimensions whose gradients point in the same directions and reduces updates for dimensions whose gradients change directions. As a result, we gain faster convergence and reduced oscillation.view on githubself.v[\"dW\"+str(i)] = ((c.BETA1*self.v[\"dW\" + str(i)]) +(eta*np.array(self.gradients[i]) ))self.v[\"db\"+str(i)] = ((c.BETA1*self.v[\"db\" + str(i)]) +(eta*np.array(self.bias_gradients[i]) ))weight_col -= self.v[\"dW\" + str(i)]self.bias[i] -= self.v[\"db\" + str(i)]Nesterov accelerated gradient (NAG)sourceHowever, a ball that rolls down a hill, blindly following the slope, is highly unsatisfactory. We'd like to have a smarter ball, a ball that has a notion of where it is going so that it knows to slow down before the hill slopes up again.Nesterov accelerated gradient (NAG) is a way to give our momentum term this kind of prescience. We know that we will use our momentum term \\(\\beta_1 v_{t-1}\\) to move the weights and biases \\(\\theta\\). Computing \\( \\theta - \\beta_1 v_{t-1} \\) thus gives us an approximation of the next position of the weights and biases (the gradient is missing for the full update), a rough idea where our weights and biases are going to be. We can now effectively look ahead by calculating the gradient not w.r.t. to our current weights and biases \\(\\theta\\) but w.r.t. the approximate future position of our weights and biases:\\(\\begin{align}\\begin{split}v_t &amp;= \\beta_1 v_{t-1} + \\eta \\nabla_\\theta J( \\theta - \\beta_1 v_{t-1} ) \\\\\\theta &amp;= \\theta - v_t\\end{split}\\end{align}\\)Again, we set the momentum term \\(\\beta_1\\) to a value of around 0.9. While Momentum first computes the current gradient and then takes a big jump in the direction of the updated accumulated gradient, NAG first makes a big jump in the direction of the previous accumulated gradient, measures the gradient and then makes a correction, which results in the complete NAG update. This anticipatory update prevents us from going too fast and results in increased responsiveness, which has significantly increased the performance of Neural Networks on a number of tasks.Now that we are able to adapt our updates to the slope of our error function and speed up SGD in turn, we would also like to adapt our updates to each individual weight and bias to perform larger or smaller updates depending on their importance.view on githubv_prev = {\"dW\" + str(i): self.v[\"dW\" + str(i)], \"db\" + str(i): self.v[\"db\" + str(i)]}self.v[\"dW\" + str(i)] = (c.NAG_COEFF * self.v[\"dW\" + str(i)] - eta * np.array(self.gradients[i]))self.v[\"db\" + str(i)] = (c.NAG_COEFF * self.v[\"db\" + str(i)] - eta * np.array(self.bias_gradients[i]))weight_col += ((-1 * c.BETA1 * v_prev[\"dW\" + str(i)]) + (1 + c.BETA1) * self.v[\"dW\" + str(i)])self.bias[i] += ((-1 * c.BETA1 * v_prev[\"db\" + str(i)]) + (1 + c.BETA1) * self.v[\"db\" + str(i)])RMSpropsourceRMSprop is an unpublished, adaptive learning rate method proposed by Geoff Hinton in Lecture 6e of his Coursera Class.RMSprop was developed stemming from the need to resolve other method's radically diminishing learning rates.\\(\\begin{align}\\begin{split}E[\\theta^2]_t &amp;= \\beta_1 E[\\theta^2]_{t-1} + (1-\\beta_1) \\theta^2_t \\\\\\theta_{t+1} &amp;= \\theta_{t} - \\dfrac{\\eta}{\\sqrt{E[\\theta^2]_t + \\epsilon}} \\theta_{t}\\end{split}\\end{align}\\)RMSprop divides the learning rate by an exponentially decaying average of squared gradients. Hinton suggests \\(\\beta_1\\) to be set to 0.9, while a good default value for the learning rate \\(\\eta\\) is 0.001.view on githubself.s[\"dW\" + str(i)] = ((c.BETA1 * self.s[\"dW\" + str(i)]) + ((1-c.BETA1) * (np.square(np.array(self.gradients[i]))) ))self.s[\"db\" + str(i)] = ((c.BETA1 * self.s[\"db\" + str(i)]) + ((1-c.BETA1) * (np.square(np.array(self.bias_gradients[i]))) ))weight_col -= (eta * (np.array(self.gradients[i]) / (np.sqrt(self.s[\"dW\"+str(i)]+c.EPSILON))) )self.bias[i] -= (eta * (np.array(self.bias_gradients[i]) / (np.sqrt(self.s[\"db\"+str(i)]+c.EPSILON))) )Complete codeAll in all the code ended up like this:view on github@staticmethoddef cyclic_learning_rate(learning_rate, epoch): max_lr = learning_rate * c.MAX_LR_FACTOR cycle = np.floor(1 + (epoch / (2 * c.LR_STEP_SIZE)) ) x = np.abs((epoch / c.LR_STEP_SIZE) - (2 * cycle) + 1) return learning_rate + (max_lr - learning_rate) * np.maximum(0, (1 - x))def apply_gradients(self, epoch): true_epoch = epoch - c.BATCH_SIZE eta = self.learning_rate * (1 / (1 + c.DECAY_RATE * true_epoch)) if c.CLR_ON: eta = self.cyclic_learning_rate(eta, true_epoch) for i, weight_col in enumerate(self.weights): if c.OPTIMIZATION == 'vanilla': weight_col -= eta * np.array(self.gradients[i]) / c.BATCH_SIZE self.bias[i] -= eta * np.array(self.bias_gradients[i]) / c.BATCH_SIZE elif c.OPTIMIZATION == 'SGD_momentum': self.v[\"dW\"+str(i)] = ((c.BETA1 *self.v[\"dW\" + str(i)]) +(eta *np.array(self.gradients[i]) )) self.v[\"db\"+str(i)] = ((c.BETA1 *self.v[\"db\" + str(i)]) +(eta *np.array(self.bias_gradients[i]) )) weight_col -= self.v[\"dW\" + str(i)] self.bias[i] -= self.v[\"db\" + str(i)] elif c.OPTIMIZATION == 'NAG': v_prev = {\"dW\" + str(i): self.v[\"dW\" + str(i)], \"db\" + str(i): self.v[\"db\" + str(i)]} self.v[\"dW\" + str(i)] = (c.NAG_COEFF * self.v[\"dW\" + str(i)] - eta * np.array(self.gradients[i])) self.v[\"db\" + str(i)] = (c.NAG_COEFF * self.v[\"db\" + str(i)] - eta * np.array(self.bias_gradients[i])) weight_col += ((-1 * c.BETA1 * v_prev[\"dW\" + str(i)]) + (1 + c.BETA1) * self.v[\"dW\" + str(i)]) self.bias[i] += ((-1 * c.BETA1 * v_prev[\"db\" + str(i)]) + (1 + c.BETA1) * self.v[\"db\" + str(i)]) elif c.OPTIMIZATION == 'RMSProp': self.s[\"dW\" + str(i)] = ((c.BETA1 *self.s[\"dW\" + str(i)]) +((1-c.BETA1) *(np.square(np.array(self.gradients[i]))) )) self.s[\"db\" + str(i)] = ((c.BETA1 *self.s[\"db\" + str(i)]) +((1-c.BETA1) *(np.square(np.array(self.bias_gradients[i]))) )) weight_col -= (eta *(np.array(self.gradients[i]) /(np.sqrt(self.s[\"dW\"+str(i)]+c.EPSILON))) ) self.bias[i] -= (eta *(np.array(self.bias_gradients[i]) /(np.sqrt(self.s[\"db\"+str(i)]+c.EPSILON))) ) if c.OPTIMIZATION == \"ADAM\": # decaying averages of past gradients self.v[\"dW\" + str(i)] = (( c.BETA1 * self.v[\"dW\" + str(i)]) + ((1 - c.BETA1) * np.array(self.gradients[i]) )) self.v[\"db\" + str(i)] = (( c.BETA1 * self.v[\"db\" + str(i)]) + ((1 - c.BETA1) * np.array(self.bias_gradients[i]) )) # decaying averages of past squared gradients self.s[\"dW\" + str(i)] = ((c.BETA2 * self.s[\"dW\"+str(i)]) + ((1 - c.BETA2) * (np.square( np.array( self.gradients[i]))) )) self.s[\"db\" + str(i)] = ((c.BETA2 * self.s[\"db\" + str(i)]) + ((1 - c.BETA2) * (np.square( np.array( self.bias_gradients[i]))) )) if c.ADAM_BIAS_Correction: # bias-corrected first and second moment estimates self.v[\"dW\" + str(i)] = self.v[\"dW\" + str(i)] / (1 - (c.BETA1 ** true_epoch)) self.v[\"db\" + str(i)] = self.v[\"db\" + str(i)] / (1 - (c.BETA1 ** true_epoch)) self.s[\"dW\" + str(i)] = self.s[\"dW\" + str(i)] / (1 - (c.BETA2 ** true_epoch)) self.s[\"db\" + str(i)] = self.s[\"db\" + str(i)] / (1 - (c.BETA2 ** true_epoch)) # apply to weights and biases weight_col -= ((eta * (self.v[\"dW\" + str(i)] / (np.sqrt(self.s[\"dW\" + str(i)]) + c.EPSILON)))) self.bias[i] -= ((eta * (self.v[\"db\" + str(i)] / (np.sqrt(self.s[\"db\" + str(i)]) + c.EPSILON)))) self.gradient_zeros()"
    } ,
  
    {
      "title"       : "Machine Learning Library in Python from scratch",
      "category"    : "",
      "tags"        : "machine learning, coding, neural networks, python",
      "url"         : "./ML-Library-from-scratch.html",
      "date"        : "2021-03-01 02:32:20 +0800",
      "description" : "Single neuron perceptron that classifies elements learning quite quickly.",
      "content"     : "It must sound crazy that in this day and age, when we have such a myriad of amazing machine learning libraries and toolkits all open sourced, all quite well documented and easy to use, I decided to create my own ML library from scratch.Let me try to explain; I am in the process of immersing myself into the world of Machine Learning, and to do so, I want to deeply understand the basic concepts and its foundations, and I think that there is no better way to do so than by creating myself all the code for a basic neural network library from scratch. This way I can gain in depth understanding of the math that underpins the ML algorithms.Another benefit of doing this is that since I am also learning Python, the experiment brings along good exercise for me.To call it a Machine Learning Library is perhaps a bit of a stretch, since I just intended to create a multi-neuron, multi-layered perceptron.The library started very narrowly, with just the following functionality: create a neural network based on the following parameters: number of inputs size and number of hidden layers number of outputs learning rate forward propagate or predict the output values when given some inputs learn through back propagation using gradient descentI restricted the model to be sequential, and the layers to be only dense / fully connected, this means that every neuron is connected to every neuron of the following layer. Also, as a restriction, the only activation function I implemented was sigmoid:With my neural network coded, I tested it with a very basic problem, the famous XOR problem.XOR is a logical operation that cannot be solved by a single perceptron because of its linearity restriction:As you can see, when plotted in an X,Y plane, the logical operators AND and OR have a line that can clearly separate the points that are false from the ones that are true, hence a perceptron can easily learn to classify them; however, for XOR there is no single straight line that can do so, therefore a multilayer perceptron is needed for the task.For the test I created a neural network with my library:import Neural_Network as nninputs = 3hidden_layers = [2, 1]outputs = 1learning_rate = 0.03NN = nn.NeuralNetwork(inputs, hidden_layers, outputs, learning_rate)The three inputs I decided to use (after a lot of trial and error) are the X and Y coordinate of a point (between X = 0, X = 1, Y = 0 and Y = 1) and as the third input the multiplication of both X and Y. Apparently it gives the network more information, and it ends up converging much more quickly with this third input.Then there is a single hidden layer with 2 neurons and one output value, that will represent False if the value is closer to 0 or True if the value is closer to 1.Then I created the learning data, which is quite trivial for this problem, since we know very easily how to compute XOR.training_data = []for n in range(learning_rounds): x = rnd.random() y = rnd.random() training_data.append([x, y, x * y, 0 if (x &lt; 0.5 and y &lt; 0.5) or (x &gt;= 0.5 and y &gt;= 0.5) else 1])And off we go into training:for data in training_data: NN.train(data[:3].reshape(inputs), data[3:].reshape(outputs))The ML library can only train on batches of 1 (another self-imposed coding restriction), therefore only one “observation” at a time, this is why the train function accepts two parameters, one is the inputs packed in an array, and the other one is the outputs, packed as well in an array.To see the neural net in action I decided to plot the predicted results in both a 3d X,Y,Z surface plot (z being the network’s predicted value), and a scatter plot with the color of the points representing the predicted value.This was plotted in MatPlotLib, so we needed to do some housekeeping first:fig = plt.figure()fig.canvas.set_window_title('Learning XOR Algorithm')fig.set_size_inches(11, 6)axs1 = fig.add_subplot(1, 2, 1, projection='3d')axs2 = fig.add_subplot(1, 2, 2)Then we need to prepare the data to be plotted by generating X and Y values distributed between 0 and 1, and having the network calculate the Z value:x = np.linspace(0, 1, num_surface_points)y = np.linspace(0, 1, num_surface_points)x, y = np.meshgrid(x, y)z = np.array(NN.forward_propagation([x, y, x * y])).reshape(num_surface_points, num_surface_points)As you can see, the z values array is reshaped as a 2d array of shape (x,y), since this is the way Matplotlib interprets it as a surface:axs1.plot_surface(x, y, z, rstride=1, cstride=1, cmap='viridis', vmin=0, vmax=1, antialiased=True)The end result looks something like this:Then we reshape the z array as a one dimensional array to use it to color the scatter plot:z = z.reshape(num_surface_points ** 2)scatter = axs2.scatter(x, y, marker='o', s=40, c=z.astype(float), cmap='viridis', vmin=0, vmax=1)To actually see the progress while learning, I created a Matplotlib animation, and it is quite interesting to see as it learns. So my baby ML library is completed for now, but still I would like to enhance it in several ways: include multiple activation functions (ReLu, linear, Tanh, etc.) allow for multiple optimizers (Adam, RMSProp, SGD Momentum, etc.) have batch and epoch training schedules functionality save and load trained model to fileI will get to it soon…"
    } ,
  
    {
      "title"       : "Conway&#39;s Game of Life",
      "category"    : "",
      "tags"        : "coding, python",
      "url"         : "./conways-game-of-life.html",
      "date"        : "2021-02-11 03:32:20 +0800",
      "description" : "Taking on the challenge of picking up coding again through interesting small projects, this time it is the turn of Conway's Game of Life.",
      "content"     : "I&nbsp;am lately trying to take on coding again. It had always been a part of my life since my early years when I&nbsp;learned to program a Tandy Color Computer at the age of 8, the good old days.Tandy Color Computer TRS80 IIIHaving already programed in Java, C# and of course BASIC, I&nbsp;thought it would be a great idea to learn Python since I&nbsp;have great interest in data science and machine learning, and those two topics seem to have an avid community within Python coders.For one of my starter quick programming tasks, I&nbsp;decided to code Conway's Game of Life, a very simple cellular automata that basically plays itself.The game consists of a grid of n size, and within each block of the grid a cell could either be dead or alive according to these rules:If a cell has less than 2 neighbors, meaning contiguous alive cells, the cell will die of lonelinessIf a cell has more than 3 neighbors, it will die of overpopulationIf an empty block has exactly 3 contiguous alive neighbors, a new cell will be born in that spotIf an alive cell has 2 or 3 alive neighbors, it continues to liveConway’s rules for the Game of LifeTo make it more of a challenge I&nbsp;also decided to implement an \"sparse\" method of recording the game board, this means that instead of the typical 2d array representing the whole board, I&nbsp;will only record the cells which are alive. Saving a lot of memory space and processing time, while adding some spice to the challenge.The trickiest part was figuring out how to calculate which empty blocks had exactly 3 alive neighbors so that a new cell will spring to life there, this is trivial in the case of recording the whole grid, because we just iterate all over the board and find the alive neighbors of ALL&nbsp;the blocks in the grid, but in the case of only keeping the alive cells proved quite a challenge.In the end the algorithm ended up as follows:Iterate through all the alive cells and get all of their neighborsdef get_neighbors(self, cell): neighbors = [] for x in range(-1, 2, 1): for y in range(-1, 2, 1): if not (x == 0 and y == 0): if (0 &amp;lt;= (cell[0] + x) &amp;lt;= self.size_x) and (0 &amp;lt;= (cell[1] + y) &amp;lt;= self.size_y): neighbors.append((cell[0] + x, cell[1] + y)) return neighborsMark all the neighboring blocks as having +1 neighbor each time a particular cell is encountered. This way, for each neighboring alive cell the counter of the particular block will increase, and in the end it will contain the total number of live cells which are contiguous to it.def next_state(self): alive_neighbors = {} for cell in self.alive_cells: if cell not in alive_neighbors: alive_neighbors[cell] = 0 neighbors = self.get_neighbors(cell) for neighbor in neighbors: if neighbor not in alive_neighbors: alive_neighbors[neighbor] = 1 else: alive_neighbors[neighbor] += 1The trick was using a dictionary to keep the record of the blocks that have alive neighbors and the cells who are alive in the current state but have zero alive neighbors (thus will die).With the dictionary it became easy just to add cells and increase their neighbor counter each time it was encountered as a neighbor of an alive cell.Having the dictionary now filled with all the cells that have alive neighbors and how many they have, it was just a matter of applying the rules of the game:for cell in alive_neighbors: if alive_neighbors[cell] &amp;lt; 2 or alive_neighbors[cell] &gt; 3: self.alive_cells.discard(cell) elif alive_neighbors[cell] == 3: self.alive_cells.add(cell)Notice that since I am keeping an array of the coordinates of only the cells who are alive, I could apply just 3 rules, die of loneliness, die of overpopulation and become alive from reproduction (exactly 3 alive neighbors) because the ones who have 2 or 3 neighbors and are already alive, can remain alive in the next iteration.I&nbsp;found it very interesting to implement the Game of Life like this, it was quite a refreshing challenge and I am beginning to feel my coding skills ramping up again."
    } ,
  
    {
      "title"       : "Single Neuron Perceptron",
      "category"    : "",
      "tags"        : "machine learning, coding, neural networks",
      "url"         : "./single-neuron-perceptron.html",
      "date"        : "2021-01-26 03:32:20 +0800",
      "description" : "Single neuron perceptron that classifies elements learning quite quickly.",
      "content"     : "As an entry point to learning python and getting into Machine Learning, I decided to code from scratch the Hello World! of the field, a single neuron perceptron.What is a perceptron?A perceptron is the basic building block of a neural network, it can be compared to a neuron, And its conception is what detonated the vast field of Artificial Intelligence nowadays.Back in the late 1950’s, a young Frank Rosenblatt devised a very simple algorithm as a foundation to construct a machine that could learn to perform different tasks.In its essence, a perceptron is nothing more than a collection of values and rules for passing information through them, but in its simplicity lies its power.Imagine you have a ‘neuron’ and to ‘activate’ it, you pass through several input signals, each signal connects to the neuron through a synapse, once the signal is aggregated in the perceptron, it is then passed on to one or as many outputs as defined. A perceptron is but a neuron and its collection of synapses to get a signal into it and to modify a signal to pass on.In more mathematical terms, a perceptron is an array of values (let’s call them weights), and the rules to apply such values to an input signal.For instance a perceptron could get 3 different inputs as in the image, lets pretend that the inputs it receives as signal are: $x_1 = 1, \\; x_2 = 2\\; and \\; x_3 = 3$, if it’s weights are $w_1 = 0.5,\\; w_2 = 1\\; and \\; w_3 = -1$ respectively, then what the perceptron will do when the signal is received is to multiply each input value by its corresponding weight, then add them up.\\(\\begin{align}\\begin{split}\\left(x_1 * w_1\\right) + \\left(x_2 * w_2\\right) + \\left(x_3 * w_3\\right)\\end{split}\\end{align}\\)\\(\\begin{align}\\begin{split}\\left(0.5 * 1\\right) + \\left(1 * 2\\right) + \\left(-1 * 3\\right) = 0.5 + 2 - 3 = -0.5\\end{split}\\end{align}\\)Typically when this value is obtained, we need to apply an “activation” function to smooth the output, but let’s say that our activation function is linear, meaning that we keep the value as it is, then that’s it, that is the output of the perceptron, -0.5.In a practical application, the output means something, perhaps we want our perceptron to classify a set of data and if the perceptron outputs a negative number, then we know the data is of type A, and if it is a positive number then it is of type B.Once we understand this, the magic starts to happen through a process called backpropagation, where we “educate” our tiny one neuron brain to have it learn how to do its job.The magic starts to happen through a process called backpropagation, where we \"educate\" our tiny one neuron brain to have it learn how to do its job.For this we need a set of data that it is already classified, we call this a training set. This data has inputs and their corresponding correct output. So we can tell the little brain when it misses in its prediction, and by doing so, we also adjust the weights a bit in the direction where we know the perceptron committed the mistake hoping that after many iterations like this the weights will be so that most of the predictions will be correct.After the model trains successfully we can have it classify data it has never seen before, and we have a fairly high confidence that it will do so correctly.The math behind this magical property of the perceptron is called gradient descent, and is just a bit of differential calculus that helps us convert the error the brain is having into tiny nudges of value of the weights towards their optimum. This video series by 3 blue 1 brown explains it wonderfuly.My program creates a single neuron neural network tuned to guess if a point is above or below a randomly generated line and generates a visualization based on graphs to see how the neural network is learning through time.The neuron has 3 inputs and weights to calculate its output:input 1 is the X coordinate of the point,Input 2 is the y coordinate of the point,Input 3 is the bias and it is always 1Input 3 or the bias is required for lines that do not cross the origin (0,0)The Perceptron starts with weights all set to zero and learns by using 1,000 random points per each iteration.The output of the perceptron is calculated with the following activation function: if x * weight_x + y weight_y + weight_bias is positive then 1 else 0The error for each point is calculated as the expected outcome of the perceptron minus the real outcome therefore there are only 3 possible error values: Expected Calculated Error 1 -1 1 1 1 0 -1 -1 0 -1 1 -1 With every point that is learned if the error is not 0 the weights are adjusted according to:New_weight = Old_weight + error * input * learning_ratefor example: New_weight_x = Old_weight_x + error * x * learning rateA very useful parameter in all of neural networks is teh learning rate, which is basically a measure on how tiny our nudge to the weights is going to be.In this particular case, I coded the learning_rate to decrease with every iteration as follows:learning_rate = 0.01 / (iteration + 1)this is important to ensure that once the weights are nearing the optimal values the adjustment in each iteration is subsequently more subtle.In the end, the perceptron always converges into a solution and finds with great precision the line we are looking for.Perceptrons are quite a revelation in that they can resolve equations by learning, however they are very limited. By their nature they can only resolve linear equations, so their problem space is quite narrow.Nowadays the neural networks consist of combinations of many perceptrons, in many layers, and other types of “neurons”, like convolution, recurrent, etc. increasing significantly the types of problems they solve."
    } 
  
]
